GPU World, 67 Cents, and the Collapsing Cost of Machine Reasoning
How this was made Verified AI
Every Intellegix briefing is generated from that day's broadcast and run through automated checks before it publishes — with a human paged on any flag. Here is the trail for this edition.
A Hacker News user posting under the name porridgeraisin published a blog post reporting a score of 44 percent on ARC-AGI-1 — the Abstraction and Reasoning Corpus benchmark designed by François Chollet at Google to measure fluid intelligence rather than pattern memorization — at a total inference cost of 67 cents. The post has ignited substantive debate about what benchmark ceilings actually mean as AI capability costs plummet.
ARC-AGI-1 was explicitly constructed to resist the statistical interpolation at which large language models excel, requiring instead novel reasoning from sparse examples. For years the design appeared to work: early frontier LLMs scored in the single digits, and the benchmark became shorthand in AI circles for a capability threshold current systems could not cross. That consensus began fracturing when OpenAI's o3 system reportedly scored above 85 percent under high-compute settings in late 2024.
The 44-percent result at minimal cost invites at least three readings. One holds that the benchmark is easier than assumed and its hard floor has been demolished. A second, more troubling reading suggests that systematic patterns in the benchmark can be exploited without genuine reasoning — a problem for its validity as an intelligence measure. A third, structurally significant reading focuses on the cost curve itself: if near-half performance on a supposedly hard reasoning task costs 67 cents today, the economic implications of that declining curve dwarf any individual score.
GPU World, a new community-maintained atlas at gpuworld.org tracking global GPU availability, pricing, comparative benchmarks, and geographic distribution of compute capacity, launched with 270 upvotes and 144 comments. The project fills a genuine information gap by synthesizing data fragmented across retailer APIs, cloud provider pricing pages, benchmark databases, and academic papers into a single curated resource — the kind of quiet infrastructure project that becomes essential over years rather than generating press releases on launch day.