Loading the context…
Loading the context…
A defined test used to assess performance on specified tasks and settings.
A benchmark supplies tasks and a method of judging results. Evaluators need to decide what counts as success, which tools are allowed and how many attempts each system receives. Repeated runs can matter because model outputs vary. A system may perform well on familiar test questions while struggling with different real-world cases; overlapping training and test material can also distort the result.
Suppose two hypothetical systems answer the same 100 accounting questions. One gets 90 correct and the other 85, producing scores of 90% and 85%. That comparison changes if the first gets five attempts per question while the second gets one. It also says little about processing invoices safely or explaining errors to a reader, because those tasks were not tested. A useful evaluation would add those relevant cases and report the conditions.
No published articles using this entry yet.