TrueFoundry open-sources TrueForge AI agent harness
TECH

TrueFoundry open-sources TrueForge AI agent harness

32+
Signals

Strategic Overview

  • 01.
    TrueFoundry launched TrueForge on August 19, 2026, an open-source, vendor-neutral AI agent harness distributed under the MIT license on GitHub and PyPI - a runtime layer that manages the full agent execution loop: model calls, tool integration, sandboxing, approvals, and session management.
  • 02.
    TrueFoundry positions TrueForge as a self-hosted, drop-in alternative to Claude Managed Agents, claiming roughly 30% lower cost per task when both run on Claude Opus 4.8, and up to 75% lower cost when TrueForge is paired with the open model GLM-5.2 instead.
  • 03.
    The harness ships with 40+ built-in tools, sandboxed execution with scoped credentials, human-in-the-loop approvals, automatic context compaction, and support for OpenAI, Anthropic, Google Gemini, GLM-5.2, other catalog providers, or any OpenAI-compatible endpoint.
  • 04.
    The harness itself is free to run; users pay only for the model calls they route through it, with an optional hosted version and AI Gateway adding centralized governance such as rate limits, budget caps, RBAC, and auditing.

The Harness, Not the Model, Sets the Bill

TrueFoundry's core pitch is that agent cost is mostly a plumbing problem, not a model-pricing problem. Its benchmark used 14 cross-system enterprise tasks drawn from DevRev's Enterprise-Bench, run across three MCP tool servers (CRM, project tracker, document store), with a fresh session per task and blind LLM-judge grading [1]. On identical Opus 4.8 traffic, TrueForge came in at $8.50 per run against $11.80 per run on Claude Managed Agents - roughly 30% cheaper - while using 3.8 million tokens and about 40 minutes of latency versus 10 million tokens and 63 minutes for the hosted alternative [1]. A third harness tested in the same benchmark, 'deepagents', was the most expensive of the three at $21 per run and 16.5 million tokens [1]. TrueFoundry attributes the gap to fewer, better-timed model calls and context compaction instead of full conversation replay - the harness decides what the model needs to see next, rather than resending everything each turn. Accuracy across all three systems landed within a task of each other, since the underlying model caps what's achievable; the differentiation is entirely in how efficiently the harness gets there [1].

Vendor Neutrality and the GLM-5.2 Swap

The more aggressive 75% cost claim comes not from a smarter harness but from a cheaper model underneath it. Swapping GLM-5.2, an open model from Z.ai, in for Claude Opus 4.8 inside TrueForge brought the same benchmark's cost down to $2.90 per run with comparable task completion [1][2]. Because TrueForge is MIT-licensed and self-hosted, that swap is a configuration change rather than a platform migration - it supports OpenAI, Anthropic, Google Gemini, and GLM-5.2 out of the box, alongside other catalog providers and any OpenAI-compatible endpoint [6]. That is the explicit vendor-lock-in argument TrueFoundry is making: as open models close the gap with proprietary frontier models on cost, an enterprise stuck inside a single vendor's managed harness cannot capture those savings without rebuilding its agent stack, while a self-hosted, model-agnostic harness can route around whichever provider is currently cheapest.

Who Grades TrueFoundry's Own Homework?

Every headline number in this story - the 30%, the 75%, the token and latency deltas - comes from a benchmark TrueFoundry designed, ran, and published itself, using its own choice of tasks, judge, and comparison targets [1]. No independent lab has reproduced these results. Analyst commentary picked up on this asymmetry without fully resolving it: Pareekh Jain of Pareekh Consulting credited TrueForge with giving enterprises more control and less vendor lock-in, while Omdia's Lian Jye Su focused praise on the governance controls rather than the cost figures themselves, flagging value for regulated industries specifically [4]. Neither analyst is quoted independently validating the cost claims. Commentator Elvis Saravia's early-access take reframes the stakes usefully: because the harness is open source, its context-truncation logic, tool-call parsing, and retry policy can be inspected line by line rather than taken on faith - which is itself an argument for treating the benchmark as a starting point for scrutiny, not a settled result.

Early Proof Points: NetApp and Automatiq

Two named enterprise users give the launch some grounding beyond TrueFoundry's own numbers. NetApp's Robert Rubin says TrueFoundry has become the central platform through which the company's agentic apps are routed, with onboarding of new users and teams happening daily and agent time-to-production shrinking as a result [3]. Automatiq's CTO George Thomas credits the platform with consolidating agent building, running, and governance into one place with the guardrails and visibility his team needed [5]. Both testimonials predate or accompany the TrueForge launch and describe TrueFoundry's broader platform rather than isolated harness benchmarks, but they are the clearest evidence in the record that production teams are already routing real workloads through the company's infrastructure.

A Launch Still Finding Its Audience

Reaction on X has been positive and enthusiastic, with early commentary from AI-focused accounts framing the release around the same harness-level cost economics TrueFoundry is promoting - fewer model calls, cheaper models slotted in, real dollar figures cited from the benchmark. TrueFoundry's own launch videos, including a founder explainer on why the harness was open-sourced, went up on its YouTube channel around the announcement but have not yet gained independent or third-party coverage on the platform; this is a days-old launch, not yet an organically viral one. Reddit reception is more muted and mixed: the official TrueFoundry account posted the announcement in its own community, and a separate account crossposted near-identical text into general AI-agent forums soon after, but engagement so far is light rather than heavy across all of them. One technical thread asked about persistent state across concurrent subagents, a real implementation question left unanswered at the time of observation, while a crosspost into a Google Jules-focused community drew pushback questioning why the post was there at all - a sign the launch's outreach was cast wider than its audience in at least one place.

Historical Context

2021
Founded in San Francisco, initially focused on machine learning model deployment software before expanding into generative AI infrastructure.
2025-02-06
Raised a $19 million Series A led by Intel Capital, with Eniac Ventures, Peak XV's Surge, and Jump Capital, bringing total funding to $21 million.
2026-08-19
Launched TrueForge, its open-source agent harness, positioning it directly against Anthropic's Claude Managed Agents.

Power Map

Key Players
Subject

TrueFoundry open-sources TrueForge AI agent harness

TR

TrueFoundry

San Francisco enterprise AI infrastructure startup, founded 2021, that built and open-sourced TrueForge as a free complement to its paid AI Gateway/governance platform, which it says processes over 1 trillion tokens daily.

AN

Anthropic (Claude Managed Agents)

Incumbent hosted agent infrastructure provider that TrueForge is explicitly benchmarked and positioned against; TrueFoundry's cost claims are measured relative to Claude Managed Agents pricing.

NE

NetApp

Enterprise early adopter already running agentic workloads on TrueFoundry; per Sr. Director Robert Rubin, the platform has become NetApp's central routing point for agentic apps in production.

AU

Automatiq

Enterprise early adopter whose CTO George Thomas credits TrueFoundry with giving the company one place to build, run, and govern agents with the guardrails and visibility it needs.

Z.

Z.ai (GLM-5.2)

Maker of the open model GLM-5.2, which TrueForge supports as a cheaper swap-in for proprietary models like Claude Opus, central to TrueFoundry's headline 75% cost-reduction claim.

Fact Check

7 cited
  1. [1] TrueForge vs Claude Managed Agents: Benchmark
  2. [2] TrueFoundry Debuts TrueForge for Building and Managing AI Agents
  3. [3] TrueFoundry Launches TrueForge, an Open-Source Vendor-Neutral Alternative to Claude Managed Agents at 50% Lower Cost
  4. [4] TrueFoundry debuts open source AI agent harness, claiming up to 75% lower costs
  5. [5] TrueForge - The Open-Source, Vendor-Neutral Agent Harness
  6. [6] truefoundry/trueforge
  7. [7] Introduction - TrueForge Documentation

Source Articles

Top 5

THE SIGNAL.

Analysts

Frames TrueForge as closing the gap between developer convenience and infrastructure control, drawing on lessons from building AI infrastructure at Meta: "The biggest lesson from building AI infrastructure at Meta was that control and convenience aren't opposites - with the right platform, you get both."

Nikunj Bajaj
Co-founder and CEO, TrueFoundry

Says TrueFoundry has become NetApp's central routing platform for agentic apps and has sped up time-to-production: "We are onboarding more users and teams every day, and it's fundamentally changed how quickly we can get an agent from an idea to something running at scale."

Robert Rubin
Senior Director of Platform Engineering, NetApp

Sees TrueForge as giving enterprises more control and reducing dependence on a single vendor's hosted stack.

Pareekh Jain
CEO, Pareekh Consulting

Flags potential value for regulated industries specifically, where TrueForge's integrated controls for budgets, access, and observability could matter more than raw cost savings.

Lian Jye Su
Chief Analyst, Omdia

Argues the agent harness layer deserves as much scrutiny as the model itself, since many perceived model failures are actually harness bugs - context truncation order, tool-call parsing, retry policy - that open-source code lets developers inspect and fix directly.

Elvis Saravia (@omarsar0)
AI researcher and commentator, early access tester
The Crowd

The Agent Harness Should Be Open Today, we are open sourcing TrueForge, the agent harness we use at TrueFoundry to build and run general-purpose agents. Because once agents do real work, the harness shapes context, tool use, execution, reliability, and cost. It is MIT licensed, vendor-neutral, and runs on your infrastructure. Same agent. Same Accuracy. Up to 75% lower cost. We used 14 tasks from DevRev Enterprise-Bench. TrueForge is 30% cheaper than Claude Managed Agents while keeping the model constant as Opus 4.8, reaching roughly the same answers while using about 40% of the tokens. Then we changed the model: running the same benchmark using GLM-5.2 with TrueForge, compared with Claude Managed Agents on Opus 4.8 at $11.8 per run, that is roughly 75% lower cost at a similar solve rate. Read detailed benchmarking blog here - https://www.truefoundry.com/blog/engineering/trueforge-vs-claude-managed-agents-benchmark/

@@truefoundry473

Sam Altman made the case for open-source harnesses in July. a month later, someone shipped it, and it's more efficient than most managed harnesses. here is the problem it was aimed at: a large share of your agent's token bill is the model rereading things it already read... @TrueFoundry's open-source agent harness, TrueForge, is built around both of those controls. TrueForge reached that score on close to a third of the tokens, with roughly 40% fewer trips back to the model, around 2.7x cheaper than Claude Managed Agents. Swapping in an open model made it sharper still: TrueForge with GLM-5.2 scored a little higher than either setup above, and the entire benchmark run cost about $3 at list prices.

@@akshay_pachaar453

We've open-sourced the agent harness we've been building at TrueFoundry. It's called TrueForge. Any model, your own infra, ~75% cheaper than Claude Managed Agents

@u/truefoundry3

OpenSourcing TrueForge Agent harness : Expecting feedback from community on the agent loop

@u/Upbeat_Pea89612
Broadcast
We Open-Sourced our Agent Harness TrueForge- Run AI Agents up to 80% Cheaper

We Open-Sourced our Agent Harness TrueForge- Run AI Agents up to 80% Cheaper

TrueForge - The Open-Source, Vendor-Neutral Agent Harness (Official Launch Video)

TrueForge - The Open-Source, Vendor-Neutral Agent Harness (Official Launch Video)

Why We Open-Sourced Our Agent Harness - TrueForge, Explained by Our Founders

Why We Open-Sourced Our Agent Harness - TrueForge, Explained by Our Founders