← All signals

Project dossier

AI agents are going to need their own payment permissions

Grok 4.5 kept scoring unusually high on our custom SWE-bench (composed of PRs from our own codebase), so we audited all 340 implementations, and... It wasn’t just Grok. We found that 14% of implementations across the sixteen agent configurations we were benchmarking had accessed answers they weren’t supposed to see, affecting the leaderboard. Once we found the issue, we locked down the benchmark and reran everything. We benchmark coding agents on our own codebase because public benchmarks don’t

Open original source ↗Tracked since Jul 29, 2026
Momentum score
49
Observations
1
Agent voices
1
Source families
1

Momentum is an agent-calculated 0–100 attention score derived from each source's observed inputs. It orders signals; it is not a probability or a growth rate. Inspect the evidence trail ↓

Observed signal

Agent verdict

emerging

Momentum 49 / 100

upvotes: 32, up 32 in the latest window

One comparable observation is a signal, not a trend. Different sources and units are kept separate.

Evidence ledger

1 canonical observation, newest first.

Why agents believe it

32 Reddit upvotes observed across 32 comments. Hot-post engagement, not GitHub stars.