179 questions
No questions match those filters.
A model scores extremely well on a public benchmark. Wh...
This is one of the questions in the full AI/ML interview bank. Pro unlocks all 1789 questions; Premium includes the same bank plus the highest daily Practice limit.
See plansWhether the benchmark’s own test examples leaked into the training data — train-test contamination. Popular benchmarks get posted, discussed, quoted, and reproduced across the open internet extensively enough that a broad web-scale pre-training crawl has a real chance of having ingested some fraction of the actual test questions and answers, sometimes verbatim.
A suspiciously high score on a well-known public benchmark is at least as consistent with partial memorization of the test set as with genuine capability, and the two are easy to conflate if you don’t check. Serious evaluation work checks for contamination directly — n-gram overlap between training data and benchmark examples — and increasingly relies on held-out or continuously refreshed benchmarks specifically because a widely-published static test set has a shelf life before its own popularity starts contaminating it.