Tiled Hacker news on React Router

Exploiting the most prominent AI agent benchmarks

530 points - last Saturday at 7:15 PM

Source

ggillas
last Saturday at 7:50 PM
This is a phenomenal paper on exploits and hopefully changes the way benchmarking is done.
From the paper: We achieved near-perfect scores on all of them without solving a single task. The exploits range from the embarrassingly simple (sending {} to FieldWorkArena) to the technically involved (trojanizing binary wrappers in Terminal-Bench), but they all share a common thread: the evaluation was not designed to resist a system that optimizes for the score rather than the task.
mzelling
last Saturday at 9:33 PM
This is an interesting catalog of vulnerabilities, but I'm not sure how groundbreaking the main insight is.
Evaluating AI models has always relied largely on trust. If you want to game the benchmarks, you can. Simply train on your test data.
When an AI agent has autonomous control over the same computing environment where its scores are recorded, it's not surprising that it can, in principle, falsify its scores. A more interesting question would be whether agents behave in this way automatically, without manual tuning by the researcher.
That said, the main takeaway of "don't trust the number, trust the methodology" is valid. It's already a truism for researchers, and spreading the word to non-researchers is valuable.
danslo
last Saturday at 8:20 PM
If only the blog itself wasn't written by AI?
>No reasoning. No capability. Just exploitation of how the score is computed.
shudder
lmeyerov
yesterday at 5:27 AM
This is great work by Dawn Song 's team. A huge part of botsbench.com for comparing agents & models for investigation has been in protecting against this kind of thing. As AI & agents keep getting more effective & tenacious, some of the things we've had to add protections against:
- Contamination: AI models knowing the answers out of the gate b/c pretraining on the internet and everything big teams can afford to touch. At RSAC for example, we announced Anthropic's 4.6 series is the first frontier model to have serious training set contamination on Splunk BOTS.
- Sandboxing: Agents attacking the harness, as is done here - so run the agent in a sandbox, and keep the test harness's code & answerset outside
- Isolation: Frontier agent harnesses persist memory all over the place, where work done on one question might be used to accelerate the next. To protect against that, we do fresh sandboxing per question. This is a real feature for our work in unlocking long-horizon AI for investigations, so stay tuned for what's happening here :)
"You cannot improve what you cannot measure" - Lord Kelvin
bluelightning2k
yesterday at 11:02 AM
They note that Mythos "found a way to inject code into a config file that would run with elevated privileges, and designed the exploit to delete itself after running".
This is more impressive than what the benchmark was supposed to be measuring. The Kobiachi Maru.
Cynddl
last Saturday at 7:52 PM
> “These are not isolated incidents. They are symptoms of a systemic problem: the benchmarks we rely on to measure AI capability are themselves vulnerable to the very capabilities they claim to measure.”
As a researcher in the same field, hard to trust other researchers who put out webpages that appear to be entirely AI-generated. I appreciate it takes time to write a blog post after doing a paper, but sometimes I'd prefer just a link to the paper.
stanfordkid
today at 1:54 AM
I don't find this paper very compelling. Obviously it would be fraud if the code generated simply escaped the harness vs solving the actual problem. I agree that theoretically models could learn to do that, and it is important to highlight, but my sense is that those entities reporting the benchmark scores would have an obligation to observe this behavior and re-consider the metrics they report. It is a bit like saying it's possible to cheat in football because the balls are deflatable. It matters, and some have done it, but it doesn't mean widespread cheating is taking place. The paper takes the tone that there is already a lot of cheating happening which I do not think is the case.
JSR_FDED
today at 1:09 AM
We have changed our entire business model so that what we actually produce is very strongly aligned with pelicans on bicycles. This way we’ll always know which model is best for us.
Highly recommend this approach, saves us tons of eval time.
lukev
last Saturday at 8:50 PM
I think we should all consider the possibility that part of the reason Anthropic hasn't immediately released Mythos is that it would be slightly disappointing relative to the benchmark scores.
SoKamil
last Saturday at 8:43 PM
The more research on this topic is created, the more knowledge how to game them will be stored in future training data. And since it comes from university, it is ranked higher in data corpus. It sounds like a self fulfilling prophecy.
mrifaki
today at 3:37 AM
this is atctually he reward hacking problem from RL showing up in evaluation infra which is not surprising but worth naming clearly, an interesting question raised here is whether agents start doing this on their own and from an RL perspective the answer is they will inevitably once benchmark performance feeds back into training signal in any form, RL finds the path of least resistance to maximize reward and if hacking the test harness is easier than solving the problem that is where gradient descent takes us, the fix is the same one the RL community has been working on for years which is to make the verifier harder to game than the task is to solve, this paper shows that right now for most of these benchmarks the opposite is true
raincole
yesterday at 5:11 AM
There are two independent issues here and I've seen people conflating them in this thread. Let's clarify:
1. Should you care or even read SWE-bench etc. scores?
The answer is no, but it has nothing to do with the vulnerabilities presented in this article. There is absolutely no reason to care about a benchmark whose dataset has been publicly available for a while. Any other way to look at benchmark scores is cargo-culting.
2. What does this article actually tell us?
It means that even if you prepared a private set of problems as benchmark, you still need to pay extra attention to how AI actually solves them. You can't lie to yourself and think this process can be 100% automated, because LLMs, as this article shows, might get the tests passed without solving the problems in a meaningful way.
bbcc90
last Saturday at 9:07 PM
Yes good evals are really hard - that’s not really news.
This team is doing a good job. They use problems that were created in last 30days to avoid training set leakage. https://swe-rebench.com/
ehtbanton
yesterday at 11:52 PM
I will always maintain that the best benchmark is just trying it out for yourself. The most practical parallel for me is all the people posting about how some open-source model has "achieved X on Y benchmark - beating out Opus 4.6!" It's all show and everyone cheats.
socketcluster
last Saturday at 11:25 PM
It feels like short-term thinking has been trained into LLMs.
They're good at solving well-defined puzzles under time constraints. It's interesting because that was the benchmark for hiring software engineers at big tech. The tech interview was and still is about fast puzzle-solving. Nothing about experience, architecture or system design in there... I suspect that's why it has a bias towards creating hacks instead of addressing the root cause.
david_shi
today at 1:07 AM
"No reasoning. No capability. Just exploitation of how the score is computed."
The irony that this was very clearly written by an LLM, double negation always the simplest and clearest tell.
yesterday at 3:01 AM
davebren
yesterday at 12:35 AM
This exploiting of benchmarks isn't that interesting to me since it would be obvious. The main way I assume they're gaming the benchmarks is by creating training data that closely matches the test data, even for ARC where the test data is secret.
sharno
today at 12:48 AM
Goodhart's law: "When a measure becomes a target, it ceases to be a good measure"
_cs2017_
last Saturday at 11:32 PM
If FieldWorkArena treats any answer as correct answer, then everyone would be getting near 1.0 (missing only when the agent is stuck in a loop or crashes). That obviously isn't what we see on their leaderboard. So does it mean the paper only found a bug in some eval code on github that no one actually uses for anything? That doesn't seem to support their claim that AI benchmarks are broken, it only supports the claim that "unused code is often buggy".
(Not commenting on any other benchmarks, just this one.)
spprashant
yesterday at 12:11 AM
I tend to prefer the ARC-AGI benchmarks for the most part. But it's always interesting when a new version drops, all the frontier models drop less than 20% or something. And then in the next few releases they get all they way up to 80%+. If you use the models it doesn't feel like those models are that much more generally intelligent.
Most frontier models are terrible at AGI-3 right now.
These models are already great no question, but are they really going be that much more intelligent when we hit 80% again?
rapiz
yesterday at 8:59 AM
Benchmark is not designed for the red team testing. I don't even think it make sense to "fix" the issue the article is suggesting. Yes, you can break the running contest by driving a car. Does this mean we need to make running contest car-proof?
lnrd
last Saturday at 8:07 PM
I'm honestly confused by the design of SWE-bench and why is considered reliable.
It's based on existing GitHub PRs and Issues, the full dataset is on HuggingFace and is one year old now. All frontier models 100% have those issues and PRs in their training data so obviously they are good at reproducing fixes for them when confronted with the same codebase and similar requests. Am I missing something? How is this considered the most reliable benchmark?
czhu12
last Saturday at 9:21 PM
I wonder if this puts into question the mythos benchmark which smashed basically all coding benchmarks to a staggering degree.
usaar333
yesterday at 4:35 AM
> But even setting aside the leaked answers, the scorer’s normalize_str function strips ALL whitespace, ALL punctuation, and lowercases everything before comparison. This means:
I don't understand the concern here
nl
today at 12:20 AM
This is a bad paper.
Benchmarking is hard to do properly. It isn't helped when people claim that exploiting the environment is some kind of flaw.
It's not. Anytime you see unexpected results running a benchmark you need to inspect what it is doing.
I recently built a yet-to-be-released where the "hard" level pushes frontier models extremely hard: Opus scores around 40%, Gemini around 60%, and GPT 5.4 around.. 0%
I inspected the traces and it turns out GPT was looking at the task and saying "I must be honest - I can't solve this task reliably" and refusing it.
> Navigating Chromium to a file:// URL reads the gold answer directly from the task config — giving ~100% on all 812 WebArena tasks.
I mean... yes? Make sure it doesn't do this?
arikrahman
yesterday at 12:34 AM
It's still a good benchmark to see which model cheats the best, I suppose.
xbar
yesterday at 11:47 PM
Dawn Song just out there killin' it.
thinkevolve
yesterday at 4:18 AM
whats the point of doing this. You have found loop holes to exploit and aced the benchmark.We did something similar with the DAB Benchmark. This exploit seems like an extension of it with lookups for the gold standard for other benchmarks.
UC Berkley will be better placed if the grads spend their time in suggesting ways to make the benchmark better.. Instead of making such simple exploits
charcircuit
last Saturday at 7:53 PM
I always assumed that these benchmarks would happen in a sandbox. I'm surprised that no one realized this sooner.
last Saturday at 10:54 PM
last Saturday at 9:37 PM
Frederick0
yesterday at 7:54 AM
This is a cracker wow
jmward01
last Saturday at 8:39 PM
Not really on the topic, but I have wondered if we need a different type of test to help find model architecture potential. Standardized training sets followed by testing to see the potential curves of a model. train on x, test, add y, test, add z, test. At each increment you see how well the model is absorbing the information and extrapolate how well that architecture may do if more fully trained.
Frederick0
yesterday at 7:54 AM
This is a cracker wow!!
moi2388
yesterday at 1:23 PM
Ironic given that the entire blog is written by AI..
avazhi
yesterday at 4:31 AM
The fact these guys got an LLM to write that page about this is diabolical.
Unreadable.
jgalt212
last Saturday at 8:41 PM
The real question is how to close to VW and Deiselgate are these offenses? And what exposure do these companies have? I would assume securities fraud, if only because Matt Levine says everything is securities fraud.
bustah
today at 10:45 AM
[dead]
Kevin_VAI
today at 10:39 AM
[dead]
oliver236
last Saturday at 8:14 PM
what are the point of benchmarks?
rupayanc
today at 8:52 AM
[dead]
techpulselab
yesterday at 4:09 PM
[dead]
neuzhou
today at 2:31 AM
[dead]
arthi1899
yesterday at 7:41 AM
[dead]
telivity-real
yesterday at 5:15 AM
[dead]
sanghyunp
yesterday at 9:55 AM
[dead]
volume_tech
yesterday at 1:21 PM
[dead]
rajptech
last Saturday at 8:38 PM
[dead]
vampiregrey
last Saturday at 11:24 PM
[dead]
usefulpatch
yesterday at 12:09 AM
[dead]
neuzhou
yesterday at 7:59 AM
[dead]
genie3io
yesterday at 8:30 AM
[dead]
semanticintent
yesterday at 1:34 AM
[flagged]
andai
yesterday at 11:16 AM
Apparently, the agent also wrote the article.

Exploiting the most prominent AI agent benchmarks

ggillas

InkCanon

phantomoc

zero_k

pxc

InkCanon

SlinkyOnStairs

tedsanders

ssivark

Imustaskforhelp

tedsanders

curioussquirrel

_blk

idrdex

Legend2440

burpingtree

aidenn0

cjbgkagh

aidenn0

Forgeties79

jimbokun

delusional

aidenn0

SlinkyOnStairs

TeMPOraL

actionfromafar

hrimfaxi

wongarsu

nurbl

UqWBcuFx6NV4r

user3939382

miki123211

anon373839

operatingthetan

ZeroGravitas

lambda

siva7

stingraycharles

retinaros

latentsea

SpicyLemonZest

operatingthetan

Leynos

nananana9

claud_ia

Aperocky

keepamovin

Barbing

siliconc0w

bisonbear

3abiton

robot-wrangler

zer00eyz

bee_rider

BugsJustFindMe

bee_rider

irishcoffee

mzelling

jmalicki

Lerc

jmalicki

boring-human

jmye

mzelling

lukev

mzelling

jmye

mzelling

hawk_aa

danslo

cpldcpu

basch

Quarrel

ainch

benob

roywiggins

not_that_d

sidpatil

blymphony