Why Most AI Benchmarks Are Misleading
7
benchmarks dont measure what matters for real use cases. heres why and what to look at instead
also worth noting – multimodal models are where the most interesting work is happening
am i the only one who thinks this?
6 replies
I actually wrote something about this a few weeks ago. the key insight for me was arxiv summaries are the only way i stay current anymore. changed how i think about the whole thing