reliability · T1 · grounding
Not in context passed
rel_ground_002
Score
100%
Turns
1
1 useful · 0 useless
Interventions
0
harness hints (HHI)
Tool calls
0/0
matched expected trajectory
Errors
0/0
recovered / total
Latency
4.08s
Tokens
267
179 reasoning
Cost
$0.00298
Trajectory timeline
Turn-by-turn replay · useless turns annotated with the classifier reason
- #0modelusefulresponse
Prompt / output inspector
instruction
Using ONLY the context, who approved release 2.4.0? If the context does not say, reply exactly 'not stated'.
context (120 chars)
Release notes — Kestrel 2.4.0 shipped on 12 May with three known issues (two cosmetic, one performance). 2.3.9 had five.
final output
not stated
expected
{"type":"abstention","shouldAbstain":true}Spans
1 recorded · tokens & cost per span
- generate4.08s85 in · 0 cached · 3 out · 179 reasoning · $0.00298 · ttft 707ms
Events
- task_started{"trial":0,"taskId":"rel_ground_002","scaffold":null}
- task_finished{"score":1,"passed":true,"taskId":"rel_ground_002","failureLabel":null}
Scorer metadata
{
"correct": true,
"toolLog": [],
"scaffold": null,
"abstained": true
}Reliabilityabstentiondifficulty 2/5