Will it run?
Research

Jerry Tworek shortens timeline to two years: researchers become remnants in AI research

By Ilse Brandt Clawpit staff
Jerry Tworek shortens timeline to two years: researchers become remnants in AI research

Jerry Tworek, who led the development of o1 and o3 and built the original codebase during his seven years at OpenAI, won't wait for anyone. In a recent interview he describes a countdown atmosphere in the lab corridors: researchers tell each other they have “a few work days left” and recommend working as long as possible. He estimates that within two years humans will cease to be a significant factor in the research process itself, not in deployment, not in oversight, but in the core development of the next models.

Tworek left OpenAI in January and founded CoreAutoAI, a startup focused on end-to-end AI-research automation. His résumé makes the forecast hard to dismiss: he signed two of the models that defined the current reasoning generation and the tool that turned code writing into a standard language-model task. He now bets that the current paradigm, the transformer architecture, has reached the edge of its capability and that the next breakthrough will come from a completely different architecture.

According to Tworek, the problem is not transformer performance but the incentive structure it imposes: scaling compute and data as a substitute for architectural depth. He reveals that during his tenure at OpenAI there were three or four serious attempts to develop an alternative architecture; none matured into a product. His conclusion is that large labs are locked onto a single path because the cost of deviating is measured in months of wasted training and massive infrastructure costs, a gamble few can afford.

The sharpest claim in the interview concerns concentration: “If you’re not the lab with the largest computational footprint, you die.” Tworek estimates that only about ten companies have ever had a realistic chance to take Anthropic’s place in the frontier race. The remaining players, including open-source labs and “neo clouds” that add a value layer on top of base models, play a different game, limited by hardware access. The number he pins on those who understand an end-to-end frontier model—30 to 50 people worldwide—demonstrates how concentrated the knowledge is.

Tworek revisits formative moments: the all-hands meeting of Ilya Sutskever in 2019 where the roadmap was almost fully realized, and the moment Jakub Pachocki handed him the keys to the GPU pool that enabled training o1. He notes with pride OpenAI’s launch of GPT-4o and describes the alignment problems of earlier generations as solved “surprisingly well”, while cautioning that the real test lies ahead.

The division Tworek now draws is clear: agents execute, write code, run experiments, and analyze results, while humans supply the creative insight, choose the direction, and ask “what’s next”. He stresses that creative writing lags behind coding not because of a research limitation but because of an absent clear reward function; in binary test code it remains subjective. When that bottleneck opens, two years from now by his clock, the loop will close and, in his view, the company itself will become a product.