• Nurse_Robot
    link
    fedilink
    English
    arrow-up
    6
    ·
    2 days ago

    If it does train, that doesn’t mean it will use that data immediately or consistently reference one data point. I don’t think your test proves or disproves anything

    • Blackmist@feddit.uk
      link
      fedilink
      English
      arrow-up
      3
      ·
      2 days ago

      Probably. The document is specific to my software, and has been there a long time. Long enough that it should be in the training data, although getting out to pull that particular bit out could be a pain.

      Maybe they only trained on documents with certain privacy options set, or use some heuristics to determine if they should use it or not. Maybe the person in the article had his data hacked and shared on a fan site.

      Hard to tell, and it’s not like Google are going to tell us either way.