𝙄 𝙗𝙪𝙞𝙡𝙩 𝙖 𝙩𝙖𝙨𝙠 𝙙𝙚𝙨𝙞𝙜𝙣𝙚𝙙 𝙩𝙤 𝙩𝙧𝙞𝙥 𝙪𝙥 𝘼𝙄 𝙘𝙤𝙙𝙞𝙣𝙜 𝙖𝙜𝙚𝙣𝙩𝙨. 𝙄𝙩 𝙜𝙤𝙩 𝙩𝙧𝙞𝙥𝙥𝙚𝙙 𝙪𝙥 𝙗𝙮 𝙩𝙝𝙚 𝙬𝙧𝙤𝙣𝙜 𝙥𝙖𝙧𝙩. Tried something new this week — contributed a task to Terminal-Bench, a benchmark used to test how good AI coding agents actually are at real terminal/computer work. Quick context on what a "task" even means here: it's basically a tiny, realistic exam question for an AI agent. An automatic checker decides if it actually got it right. The task I built: a video had been sped up for broadcast and had a few extra seconds tacked onto the front, so the subtitles no longer lined up with it. The agent had to figure out, purely from clues buried in the video files, exactly how much faster it was playing and exactly how much extra time had been added — then use that to shift every subtitle line back into sync. 𝙏𝙝𝙚 𝙚𝙫𝙖𝙡𝙪𝙖𝙩𝙞𝙤𝙣 𝙝𝙖𝙧𝙣𝙚𝙨𝙨 𝙩𝙚𝙨𝙩𝙚𝙙 𝙖𝙜𝙖𝙞𝙣𝙨𝙩 𝘾𝙡𝙖𝙪𝙙𝙚 𝙊𝙥𝙪𝙨 4.8, 𝙩𝙝𝙚 𝙢𝙤𝙙𝙚𝙡 𝙨𝙤𝙡𝙫𝙚𝙙 𝙞𝙩 𝙘𝙤𝙧𝙧𝙚𝙘𝙩𝙡𝙮 𝙤𝙣𝙡𝙮 𝙤𝙣𝙘𝙚 𝙤𝙪𝙩 𝙤𝙛 𝙛𝙞𝙫𝙚 𝙞𝙣𝙙𝙚𝙥𝙚𝙣𝙙𝙚𝙣𝙩 𝙖𝙩𝙩𝙚𝙢𝙥𝙩𝙨. Here's the part I found genuinely interesting: it nailed the hard bit every single time — correctly working out the speed change from the file data, which is the part I'd actually built the task around. What got it, most of the time, was something much smaller. The clues for the "extra time added" part had a tiny bit of expected noise in them — completely normal, still clearly pointing to the same answer. But instead of accepting that and moving on, the model treated the noise as a red flag, decided something must be wrong, and burned all its time trying to dig up an "exact" answer that didn't actually exist. It never got around to answering the actual question. That lines up with a pattern I think is broader than this one task: LLMs tend to be genuinely strong at deep, structured reasoning — the kind of problem with one right path. Where they struggle is judgment under uncertainty: knowing when a slightly messy, real-world signal is good enough to act on, versus when it's actually a problem worth chasing. Models, at least right now, seem to default to over-verifying instead of trusting a reasonable answer and moving forward — and that's often more costly than just being wrong.
#TerminalBench #AIAgents #LLMEvaluation