Inside Project Mercury: How OpenAI Built a Banker Before It Built a Product
A year before this launch, OpenAI ran a secretive internal effort called Project Mercury, paying over 100 former investment bankers from firms like JPMorgan, Goldman Sachs, and Morgan Stanley roughly $150 an hour to build industry-standard Excel financial models [1][2][3]. It was, in effect, manufacturing the training data an AI would need to think like a junior banker before any product existed to sell. That investment shows up in the benchmark numbers OpenAI is now citing: on OfficeQA Pro, a test built from 133 questions drawn from roughly 89,000 pages of Treasury Bulletin filings, the new GPT-6 Astra model scored 69.9% correctness, ahead of OpenAI's own GPT-5.6 Sol at 60.2% and Anthropic's Claude at 62.4% [4]. OpenAI also reported sharp drops in data-connector error rates across several third-party providers - Daloopa's error rate fell from 7.53% to 2.57%, and S&P Global's from 6.84% to 2.66% [4]. The Project Mercury story matters because it reframes this launch: it is not a clever prompt wrapped around a general model, but the payoff of a deliberate, expensive data-acquisition campaign built specifically to reproduce the judgment of the people it may ultimately displace.



