BenchMIRT: What are LLM benchmarks actually measuring?
Understanding What Benchmarks Really Measure
Evaluating large language models requires standardized benchmarks, but a critical question underlies this practice: what are these benchmarks actually measuring? A new analytical framework called BenchMIRT addresses this fundamental concern by applying Multi-dimensional Item Response Theory (MIRT) to examine the structure and validity of LLM