So funny you should mention that, I worked at a company that dealt with Linux and third-party proprietary code. They kept the software developers highly segregated because they feared accidental copyright infringement. They thought at the time that even a human learning and accidentally reproducing something they remembered from working on proprietary code too risky.
The practical difference is that the third-parties were likely to sue, but the corpus of trained data is pretty much open source projects that may have a right to sue but in practice everyone knows they aren’t going to be able to chase down violations due to lack of resource.
Whole thing is a nightmare and I hate it. Luckily nothing I do for work is open source so even if my code accidentally resembles some other codebase I’ve worked on, nobody will find out lol
So funny you should mention that, I worked at a company that dealt with Linux and third-party proprietary code. They kept the software developers highly segregated because they feared accidental copyright infringement. They thought at the time that even a human learning and accidentally reproducing something they remembered from working on proprietary code too risky.
The practical difference is that the third-parties were likely to sue, but the corpus of trained data is pretty much open source projects that may have a right to sue but in practice everyone knows they aren’t going to be able to chase down violations due to lack of resource.
Whole thing is a nightmare and I hate it. Luckily nothing I do for work is open source so even if my code accidentally resembles some other codebase I’ve worked on, nobody will find out lol