Sry for not answering for a while but I am not sure myself on a lot of this.
I know for a while the big companies scraped petabytes of data into datasets for training. It was quantitative, the more the better.
About 2 years ago people found out removing bad data is really helpful and more is not necessarily better. Datasets got smaller and more focused on beneficial data, which of course looks very different depending on the purpose of the model.
Basically what happens is someone trains a model, then tests the model, then trains then tests, all the while they are slowly tweaking stuff and rewriting basic parts of how it’s trained or how the model does inference “using the model”. And they try to see how to make it so the model performs the best. It’s like 3 fields of science combined and a lot of the time they release a whole scientific paper together with the model. Especially if there’s some new technical discovery they stumbled upon in training which might advance the field.
But I can’t do that, I am smart and experienced in that field yet still I will happily admit I could not keep up with those brilliant scientists under any circumstance.
When I wanna run a local AI, I download a model. For consumers you expect it to be a maximum of 30gb for the model and maybe another 5gb for libraries and stuff. Then using those libraries you run the model, feed it your data (the microphone input) and let it have at it. You rarely do any more training, because you will almost never do it better than a scientist who dedicates most of their time to that specialty.
Think of the model itself as a chaotic big blob. Like a zip archive but inside is a bunch of unreadable data, usually called “weights”. Those weights give the AI a way to calculate an output from an input. It’s not a dataset, but it’s everything that the model could learn from the dataset when it was trained. And then I myself never have to deal with terabytes of data from a dataset, it’s all neatly packed into a 30gb model. And I can just run that.
I can explain it in more detail but even I can’t grasp the full technical details anymore because every step has tons of optimizations and transformations baked in, but the very basic model still functions as we are used to from deep neural networks. If you are interested in learning, that’s a good way to start.
Anyway, you sly dog caught me monologuing. Thank you for letting me share all of this stuff :)
No, not a problem at all :) thanks for getting back to me.
I didn’t realize you couldn’t just feed it raw data, but that makes sense. That’s too bad, I would have gone all in on local ai if it was something end users could specifically curate their individual models with. Definitely still sounds useful though.
30gb is still way smaller than what I was expecting. That’s pretty actionable, just like uninstall a single videogame.
Thank you, and I hope you have a great day as well
Yeah you can use your own data but it’s extremely unlikely that you get comparable results and it takes much more time and more trial and error to train such a model yourself.
Even just using a model that’s trained on royalty free works like the dolphin dataset is more feasible than using your own data I assume.
But yes, legally and ethically the boundaries of personal identity and copyright will be part of the discussion for at least a few more years.
Sry for not answering for a while but I am not sure myself on a lot of this.
I know for a while the big companies scraped petabytes of data into datasets for training. It was quantitative, the more the better.
About 2 years ago people found out removing bad data is really helpful and more is not necessarily better. Datasets got smaller and more focused on beneficial data, which of course looks very different depending on the purpose of the model.
Basically what happens is someone trains a model, then tests the model, then trains then tests, all the while they are slowly tweaking stuff and rewriting basic parts of how it’s trained or how the model does inference “using the model”. And they try to see how to make it so the model performs the best. It’s like 3 fields of science combined and a lot of the time they release a whole scientific paper together with the model. Especially if there’s some new technical discovery they stumbled upon in training which might advance the field.
But I can’t do that, I am smart and experienced in that field yet still I will happily admit I could not keep up with those brilliant scientists under any circumstance.
When I wanna run a local AI, I download a model. For consumers you expect it to be a maximum of 30gb for the model and maybe another 5gb for libraries and stuff. Then using those libraries you run the model, feed it your data (the microphone input) and let it have at it. You rarely do any more training, because you will almost never do it better than a scientist who dedicates most of their time to that specialty.
Think of the model itself as a chaotic big blob. Like a zip archive but inside is a bunch of unreadable data, usually called “weights”. Those weights give the AI a way to calculate an output from an input. It’s not a dataset, but it’s everything that the model could learn from the dataset when it was trained. And then I myself never have to deal with terabytes of data from a dataset, it’s all neatly packed into a 30gb model. And I can just run that.
I can explain it in more detail but even I can’t grasp the full technical details anymore because every step has tons of optimizations and transformations baked in, but the very basic model still functions as we are used to from deep neural networks. If you are interested in learning, that’s a good way to start.
Anyway, you sly dog caught me monologuing. Thank you for letting me share all of this stuff :)
Hope you have great day ^^
No, not a problem at all :) thanks for getting back to me.
I didn’t realize you couldn’t just feed it raw data, but that makes sense. That’s too bad, I would have gone all in on local ai if it was something end users could specifically curate their individual models with. Definitely still sounds useful though.
30gb is still way smaller than what I was expecting. That’s pretty actionable, just like uninstall a single videogame.
Thank you, and I hope you have a great day as well
Yeah you can use your own data but it’s extremely unlikely that you get comparable results and it takes much more time and more trial and error to train such a model yourself.
Even just using a model that’s trained on royalty free works like the dolphin dataset is more feasible than using your own data I assume.
But yes, legally and ethically the boundaries of personal identity and copyright will be part of the discussion for at least a few more years.