Thinking that harnesses and agents can improve without human involvement, even if it works, will yield a much lower growth rate than if a human gets involved in the loop.
This is the best take. Return on capital is diminishing from intense competition so improvements get hidden away like the gems they are.
1) Frontier labs have no incentive to give the general public their best anymore; it's instantly distilled off of them. Why not charge governments and big corps real money to use the real good stuff instead?
2) So we get distilled-off-frontier public APIs like 5.6 and Fable/Opus. And the open source labs are distilling off of those.
3) There's a lot of benchmark hacking right now among all the publicly available models, actual usability of Opus for coding is far below its benchmarks suggest.
4) But context window, cybersecurity, logical coherence, and tool usage are absolutely better on Fable and Sol. It looks to me their internal tools definitely even better and not plateauing. But we won't get to use it.
i learned that no matter how good i am at writing prose, it doesnt matter if nobody reads it. so i write for AI agents instead, hoping that it'll get picked up by an agent and find propagation that way
You can actually run these benches yourself, as Frontier-Bench is open source.
Also have a look at these other coding benchmarks I audited.
Frontier-Bench v0.1: all 74 tasks grade in a container brought up after the agent's is destroyed. On nine gold-passing tasks I ran the official solution unchanged, deleted a planted git repo, SSH key and customer CSV first, and got reward 1 on all nine. https://june.kim/auditing-frontier-bench
Terminal-Bench 2.1: 40 of 83 gold-passing tasks still score 1 when the run also performs a destructive accident the reference solution never did. https://june.kim/terminal-bench-frame
DeepSWE: 1 of 113 published gold patches breaks its own tests, and the per-task verdicts behind the score aren't retrievable. https://june.kim/auditing-deepswe-v1-1
MirrorCode: better built than most, with 2 of 25 targets reachable from published specs rather than from the artifact it hands you. https://june.kim/auditing-mirrorcode
I'd buy that if lots of high quality PRs were sitting around waiting to be merged, but they don't exist. There's barely any evidence that AI is even real in the open source world, other than the slop of course
Man, truly absolutely gross to see this, this is some incredibly unethical engineering here. This is literally why people hate AI slop. Its not surprising to me that you didn't want to even write the post yourself
>It cost me 53 merges, 63 closures, a billion Opus tokens, and eight account blocks
I can't believe people think this is acceptable. You didn't fix any problems, you just managed to deliberately sneak slop past maintainers without their knowledge. There's no review by you here that any of the PRs actually fixed anything (presumably because software engineering is a learned skill), you just assume that the AI fixed the bugs without checking any of it apparently because.. you feel you're important?
the issues were already validated by the maintainers, and the maintainers merged it into their repo voluntarily. When maintainers accepted the PRs, they were the ones who found it acceptable. I found that well-tested PRs are more likely to get merged, so I deliberately had it pick bugs that were easy to verify. I didn't need to get involved after I specified that criteria.
Humans don't get a 100% merge rate either, so it's a matter of degree to which stranger is more effective at contributing to open source. It was an experiment where I accumulated lessons about how OSS works, and the lessons are open and available for your viewing here:
https://github.com/kimjune01/sweep/blob/main/HYPOTHESIS_GRAP...
AI is good at fixing a narrow subset of issues/tasks in GitHub, so humans still need to be involved in most tasks.
reply