I can name groups of people I interact with who lean both ways.
It’s still a commonly held belief that “Facebook sells your data” and it’s cool to be cynical about everything tech in many social scenes. Conceding that a tech company might be honest about something will get you classified as a bootlicker depending on who you talk to so the only winning move is to be super cynical.
Among actual professionals I work with in tech and legal, almost nobody holds a belief that these companies are blatantly lying to their customers (and zero of their employees are whistleblowing it, while said companies also have employees trying to whistleblow AI safety on Twitter daily)
Yes, I am talking about working professionals who use LLMs. Before this thread, I would have considered it surprisingly and singularly naïve if someone told me they trusted OpenAI. I still believe the common and correct take is that these companies are largely training on customer data against their consent.
I don't think they are "blatantly" lying either, just normal bog-standard lying that we've all come to accept. It's a profitable and competitive tech company.
We have already seen this lying. The toggles are opt-out, not opt-in. When you sign up, you agree to binding arbitration, which is effective for preventing lawsuits in the US. The toggles are regularly turned back on without our consent on ChatGPT and Claude. OpenAI's "don't train on my content" setting isn't even in the ChatGPT interface.
As far as I know, they haven't suffered even a tiny controversy in public opinion over any of this at all.
There's nothing to whistleblow about when it's public knowledge.
How many of the people who checked those boxes have cryptographic proof they did it? How many of those people have opted out of the arbitration clause? How many of those people would be able to claim damages? Would the amount of people who satisfy all three questions be large enough to make it worth _not_ training on user data?
I also don’t think these companies are lying at all, but I definitely think they’re training on all your data, toggle or not.
It’s truly trivial to “anonymize” and distill your prompts and model output. They could use just about any off-the-shelf cheap model for this. In fact, their TOS explicitly allows this, even with the toggle checked.
What that probably means is that the EXACT content of your prompt is secret. But the actual ideas are not. If you discover something truly novel, then yeah they get that. They can absorb trends in consumer behavior, too.
I’m sure if someone had access to all my paraphrased prompts, which retain 0% of my exact wording, they could find out literally everything about me. It’s a bit like how collecting metadata is as good (or better!) than collecting the real data.
And we all know “anonymizing” data doesn’t really exist like we think it does. Just removing names and identifiers doesn’t make anything anonymous for motivated actors. Or… say… an LLM that is trained to recognize patterns in text. Which is, like, all of them.
It’s still a commonly held belief that “Facebook sells your data” and it’s cool to be cynical about everything tech in many social scenes. Conceding that a tech company might be honest about something will get you classified as a bootlicker depending on who you talk to so the only winning move is to be super cynical.
Among actual professionals I work with in tech and legal, almost nobody holds a belief that these companies are blatantly lying to their customers (and zero of their employees are whistleblowing it, while said companies also have employees trying to whistleblow AI safety on Twitter daily)