This essay was written with Barath Raghavan, and originally appeared in IEEE Spectrum. Major benchmarks measure what AI can do. None measure whether it does what you mean: the distance between what you ask an AI to do and the unspoken assumptions about how you want the AI to do it. We propose a new metric: the Genie coefficient. There’s often a gap between one person’s request and another’s understanding. Most of the time, we bridge it using general knowledge. For example, if you ask a friend to get you coffee, they’ll pour a cup from the pot or buy one from a coffee shop. They won’t bring you a bag of raw beans or snatch a cup from a stranger and hand it to you. You never specified any of this. You never had to...
Much of AI is a scam. Probably the large majority.
I’m not sure you’ve quite captured how overfitting functions in current language models. Nor how baselines and benchmarks are used for machine learning products. For serious models of the current year, overfitting is something that happens when you train too long on the same data. It happens less if you add new data, and has nothing to do with new benchmarks and evaluations. You can read about overfitting more generally (mostly the classical kind) on wikipedia.
I think Bruce Schneier has a reasonable understanding of how benchmarks, public policy, large language models, and overfitting are at work here.
Everything LLM is a scam and a large portion of other neural network tech has also proven unreliable.
That’s one way for everything you say to lose credibility.
OK, Slopper.
Ah, confirmed bias
Would you please cite your sources? I’d like to read more about this topic.
While I’m at it I’ll show you the proof that god definitively does not exist. /s
OK buddy. It’s been running your auto correct, search, a bunch of medical scans, and map routing for awhile now with no complaints. But now that we’ve learned some LLM vocabulary…
lmfao
so far off the mark with that one.