Your model can code, so what?
In a new era of harness engineering where cheap models can code "like" humans, what is stopping Artificial General Intelligence?
We measure progress toward AGI and superintelligence with transparent, citable evidence: capability trends, benchmark saturation, and the open gaps between the systems we have and the systems we’re trying to build.
An independent research institute measuring progress toward artificial general intelligence and superintelligence — and assessing whether today’s methods can get us there.
The paper that turned capability into an engineering forecast: loss falls as a smooth power law in compute, data, and parameters. The empirical backbone of the "scaling is enough" thesis we continually stress-test.
The "Chinchilla" result recalibrated the scaling recipe: most large models were badly undertrained on data. It reset everyone's compute-to-data ratios and remains essential for any honest forecast of where the curves go next.
A sprawling, contested case that GPT-4 shows early signs of generality. Read it alongside its critics: it is as much a document about how hard AGI is to evaluate as it is a claim about the model.