Tools Bench.

Product launches and open-source repos with enough signal to earn a second look.

Last Brew Time: Sep 4, 2026, 10:52 AM PT

Insight

This run's builders are spending less effort shipping new apps and more on disciplining the agent's own habits

Featured

GitHub203.9K

The open source coding agent.

Market Signal

Why It Has Market Pull

opencode is one of the most prominent open-source coding agents available, with a scale of adoption and public debate that puts it among the top tier of AI developer tools. It is backed by an established team (formerly SST, now operating as Anomaly) and has weathered a high-profile dispute with a major AI lab without losing momentum.

  • 204,019 GitHub stars and 26,626 forks, more stars than Claude Code by some public counts
  • A Hacker News thread about a legal dispute with Anthropic reached roughly 1,099 points and 546 comments
  • Grew from about 172,000 stars in June 2026 to over 204,000 by September 2026
  • Now supports ChatGPT-subscription login natively following a partnership with OpenAI
  • Built by the team behind SST, rebranded as Anomaly, with a long history of shipping developer tools

feedbacks

What People Are Saying

  • "Very similar to Claude Code in terms of capability"independent guide

  • "Using opencode with Anthropic OAuth violates ToS and results in a ban"GitHub issue

  • "Developers accused Anthropic of anti-competitive behavior after the login change"HN comment

  • "Forks reverting the change appeared within hours, alongside community plugins restoring functionality"developer community

  • "Connects to 75+ AI providers and gives full control over which models process your code"project documentation

  • "One of the most-starred open-source coding agents available today"independent guide

GitHub174.0K

Public repository for Agent Skills

Market Signal

Why It Has Market Pull

This is Anthropic's own official repository defining the Agent Skills format, not an independent project — it already anchors a cross-industry standard. Adoption spans major coding tools beyond Claude itself, making this one of the strongest signals in this batch.

  • 174,000+ GitHub stars and 20,600+ forks, roughly doubling since February 2026
  • First-party repository maintained directly by Anthropic with commits landing almost daily
  • Format adopted by roughly 40 client tools per an industry directory, including GitHub Copilot, Cursor, VS Code, and Gemini CLI
  • 1,208 open issues and 1,119 watchers indicate a large, engaged developer base
  • Enterprise admin controls and a partner skills directory were added as the project matured

feedbacks

What People Are Saying

  • "Doubled in stars since February 2026 as more tools adopted the underlying spec"Dev.to article

  • "Now treated as the reference implementation other agent platforms build against"Industry directory

  • "Documented as natively supported by at least one major coding assistant"Vendor documentation

  • "1,200+ open issues reflect heavy day-to-day use rather than a dormant spec repo"GitHub issue tracker

  • "Regular substantive content updates to individual skills rather than just churn"GitHub commit history

  • "Cited as the basis for a growing partner skills directory"Vendor blog

GitHub31.0K

TimesFM (Time Series Foundation Model) is a pretrained time-series foundation model developed by Google Research for time-series forecasting.

Market Signal

Why It Has Market Pull

TimesFM is Google Research's zero-shot time-series forecasting foundation model, and it has moved well past a research curiosity into production infrastructure — it's built into BigQuery ML (generally available) and adopted in Databricks' forecasting tooling. With a peer-reviewed ICML paper, a benchmark-leading successor (TimesFM-3), and steady releases through August 2026, this is a credible, actively maintained product with real enterprise reach.

  • 31,000 GitHub stars and nearly 3,000 forks, verified via the GitHub API, still active with a v3.0.0 release on August 28, 2026
  • Ranks #1 on the GIFT-Eval benchmark across 28 datasets with just 200M parameters
  • Generally available inside Google's BigQuery ML, plus Google Sheets and Databricks integrations
  • Originally presented at ICML 2024, giving it independent academic vetting
  • Latest version (TimesFM-3) trained on over a trillion real-world and synthetic time points, reported as state-of-the-art for multivariate forecasting

feedbacks

What People Are Saying

  • "TimesFM 2.5 ranks #1 on GIFT-Eval across 28 datasets with just 200M parameters"Research blog

  • "The boom of foundation models in time series forecasting"Dev.to article

  • "Google's new forecasting model beats everyone — you can't use it at work (yet)"Tech news site

  • "Now generally available inside BigQuery ML for production forecasting workloads"Product documentation

  • "Adopted for forecasting across retail, finance, observability, manufacturing, healthcare, and natural sciences since its 2024 debut"Industry report

  • "Outperforms most statistical baselines like ARIMA and ETS, and matches or beats several deep learning models trained specifically on the target series"Research report

Product Hunt329

Describe any workflow in plain English. Agent Builder compiles it into a coded automation that runs like software. When a run breaks, Airtop automatically investigates, rebuilds the step, and verifies the fix on a real test run. Airtop Agents run securely in the cloud and are up to 100x more efficient than traditional LLM-per-step AI Agents.

Market Signal

Why It Has Market Pull

Airtop is an established, well-capitalized company launching a genuinely new capability, which is a stronger signal than most fresh launches. It has raised a $25M Series A from notable investors and reportedly does $1.7M in revenue with a lean 15-person team, and this self-healing agent feature drew real, substantive product feedback rather than generic praise.

  • Raised a $25M Series A with investors including Andreessen Horowitz, Bessemer Venture Partners, Aaron Levie, and Greg Brockman
  • Reported $1.7M in revenue with a 15-person team
  • 329 upvotes on its Product Hunt launch
  • Reviewers describe it as a strong alternative to established browser-automation tools
  • Specific, workflow-level feedback from users trying it on real automations (not just launch-day praise)

feedbacks

What People Are Saying

  • "Self-healing is great, but you still want to understand what changed on the page."Product Hunt comment

  • "One workflow failed because of a region-locked, US-only proxy; selectable proxy regions are reportedly planned."Product Hunt comment

  • "Praised for smooth onboarding, UX, and integrations, notably with Make."Product Hunt review

  • "Seen as a strong alternative to Browserbase-style automation tools."Product Hunt comment

  • "$1.7M in revenue with a 15-person team."Industry profile

  • "$25M Series A backed by well-known venture and angel investors."Funding database

Product Hunt221

Atlas is World Labs' omni world model. It takes text, images, video, and 3D, then generates camera-controlled 1440p video up to a minute, reconstructs scenes from a few photos, and simulates space-time for robotics. Early access.

Market Signal

Why It Has Market Pull

This is one of the most credible entries in the batch: a well-funded company led by a well-known AI researcher, generating substantial press coverage and third-party technical commentary. It's early access only with no public code or paper, so hands-on user validation is still limited, and some coverage flags that the company's own comparison claims aren't independently verified.

  • Company has raised roughly $1.23 billion, including a $1 billion round backed by Autodesk, AMD, NVIDIA, and Fidelity
  • Founder's launch post reached wide visibility on X and was covered by multiple tech outlets
  • Company-run blind evaluations report 75-94% preference over rival models, though these are self-reported and not yet independently reproduced
  • Currently early access with select partners only; no public weights, code, or paper
  • Independent technical write-ups flag unresolved failure modes (identity drift, thin geometry, accumulating errors on long camera paths)

feedbacks

What People Are Saying

  • "Introducing Atlas, a first-of-its-kind multimodal world model trained from scratch."X post

  • "Third-party raters preferred Atlas over rival models by wide margins in blind evaluations."Industry blog

  • "The launch language is aggressive; claims like 'pixel-perfect' are the company's own, not independently established."Industry blog

  • "Likely failure modes include invented off-camera structure and identity drift, none of which are quantified at launch."Industry blog

  • "Coverage of the launch appeared across multiple tech and finance outlets within a day."Press coverage

  • "No public weights, code, or paper have been released; access is limited to select partners."Company blog

Sources

GitHub

Skills for Real Engineers. Straight from my .agents directory.

248.2K

The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

Makes your AI agent think like the laziest senior dev in the room. The best code is the code you never wrote.

🪨 why use many token when few token do trick — Claude Code skill that cuts 65% of tokens by talking like caveman

Agent skill that removes signs of AI-generated writing from text

Product Hunt

336

Nex automates workflows like CRM clean up, large list qualification, revenue recovery, at large volumes, that break general purpose agents. Built by ex-HubSpotters who ran its three largest customer platforms and backed by HubSpot founder Dharmesh Shah.

MagiCrew is an open-source AI Agent platform that gives everyone their own AI workforce. Instead of simply chatting with AI, deploy specialized digital workers that research, analyze, create reports, generate presentations, and complete real business tasks. With multi-agent collaboration, enterprise controls, and deliverable-ready outputs, MagiCrew helps teams turn AI from a tool they use into a workforce they can manage.

263

Omi captures your screen and conversations, creates tasks, reminders and advice, helping you find anything you saw or discussed. Turn calls into summaries and action items, ask questions and get personalized responses! Omi is open source, local and works with your own AI keys, giving you full control over what gets recorded, paused, or deleted.

Tabbit is the AI browser that knows what you’re working on and can get the work done for you. Give it the pages, screenshots, and local files that matter, and Tabbit Agent can work across the web now or on a schedule. It delivers usable HTML, PDFs, and presentations, then saves the workflow as a Skill you can run again.

Higgsfield Genjutsu is an AI video-to-video tool for creators, influencers, and marketers who want to transform existing footage without a full reshoot. Its Motion Transfer and Object Swap features let you change characters, objects, outfits, locations, or styles while preserving the original motion, timing, and shot structure, making content variations faster and more flexible.

Blume watches your coding agent sessions locally and turns what it learns into better agent context. Repeated corrections become rules, workflows become skills, and your agents stop making the same mistakes. Claude code, Codex and Cursor.

YC Launch

Hacker News

We run a SaaS that handles petabytes of data. Our SRE team experimented with using claude, openclaw, langchain, etc. within our incident response workflows. We struggled with overflowing context, lethal trifecta vectors, hallucinations, and burned a lot of frontier tokens mostly on easy work. Approval fatigue was a challenge, and we drew a hard line at relaxing permissions in production. Long story short, we built and open-sourced AURA, a Rust-based harness specifically designed for the type of... (25 points, 6 comments).

I’m Nate, the founder of Ardent. We just shipped our public beta, and we’d love your thoughts! Ardent is an agent running in a desktop (Electron) app built to help with knowledge work, designed for less-technical people outside engineering. Think Codex or Claude Cowork, but built around collaboration and customization. I know, I know, it’s yet another agent harness! Ardent is a little different – it leans heavily on codegen to solve problems. Most agents are basically just a bag of tools and a w... (10 points, 2 comments).

Hi HN. We built an API context registry to help coding agents (like Claude Code) generate production-ready API integration code without blowing through token limits. We build a lot of API integrations. In our experience, most coding agents write basic client calls fine, but consistently stumble on details that make code shippable, like idempotent retries, rate-limiting and Auth token management. We tried all the existing approaches of injecting context into coding sessions: - Markdown dumps deli... (8 points, 1 comments).

Hi HN, I'm Bor Shev, a composer and developer. Over the past two years I've been developing ShevtoneAudio Orchestrator. The idea is simple: instead of generating a finished piece of music and replacing the composer, Orchestrator takes the composer's own MIDI and develops it into a full orchestration. It analyzes the musical material — harmony, melody, rhythm, dynamics, structure and orchestral density — and creates an arrangement across strings, brass, percussion and other sections. The importan... (6 points, 4 comments).

Hello HN, I'm Ali, building Decispher. The problem we're working on is that coding agents repeatedly rediscover context that already exists inside an engineering organization. A developer working on a feature can combine information from previous PRs, Jira tickets, Slack discussions, ownership boundaries, architectural decisions and their own experience. Coding agents usually start with a prompt and a repository, then spend tokens searching for that same context—or miss it entirely. Decispher is... (5 points, 0 comments).

My company stopped allowing Navicat because of compliance policy. DBeaver worked, but I don't like it. I kept running into two recurring issues. First, production investigations often require several queries across different databases. I would query one table, copy an ID into another query, wait for the result, and repeat the process several times. I started wondering if there was a way to handle this without writing a separate script for every case. What if I had a simple form where I could ent... (4 points, 1 comments).

HF Spaces

Demo of the Collection of Qwen Image Edit LoRAs Qwen-Image-Edit-2511-LoRAs-Fast is a Hugging Face Space tagged with gradio, mcp-server, region:us. It has 2714 likes on Hugging Face.

Real trained RL policies for the Microduck robot, running fully in the browser: MuJoCo compiled to WebAssembly steps the physics, onnxruntime-web runs the policy network at 50 Hz. No server, no backend. Two locomotion variants of the same robot are included: legs (walking, the default) and rollers (the wheeled skating variant). Press M (or hold D-pad up ~1 s on a gamepad) to switch; the roller model, meshes and policies are lazy-loaded on the first switch. | Mode | Checkpoint | What it does | |--------|-----------|--------------| | Run (legs) | BESTalphawalking.onnx | Velocity-tracking locomotion (arrows / WASD to steer) | | Sit | BESTalphasitstand.onnx | Sits down on its hull, stands back u...

Video generation with a synchronized soundtrack MiniMax-H3 — unquantized, split across two Spaces Joint video and soundtrack out of a single denoising pass, at bfloat16 with no quantization anywhere. This Space is the denoising half: the 61.73 GiB transformer and the two autoencoders. The 62.14 GiB Qwen3-VL conditioner runs in qwen3vl-conditioner, which this Space calls over the gradio API for every request. The weights are the public MiniMaxAI/MiniMax-H3 diffusers checkpoint. MiniMax-H3 is 195.9 GiB in bfloat16 and a ZeroGPU Space is evicted at 150 GB of storage. An unquantized single Space is therefore impossible, which is why quantized demos of it run NVFP4 or float8 weights. Cut the Mini...

Unified text-to-image and image editing model Text-to-image and image editing demo for sensenova/SenseNova-U1.5-8B-MoT, a natively unified multimodal model (18B params, bf16) built on the NEO-unify architecture. Leave the image upload empty for text-to-image generation, or upload one or more images and write an edit instruction for image editing. Advanced options exposes denoising steps, guidance scale, timestep shift, image guidance (editing) and the seed. The step count defaults to 28 rather than the model card's 50: a fixed-seed A/B found 28 keeps composition, prompt adherence and text rendering intact — losing only some micro-texture in landscape and skin, and nothing measurable when edi...

Blind A/B ranking of MiniMax-H3 acceleration variants Human-judged ranking of ~26 MiniMax-H3 acceleration variants over a 200-prompt corpus, from blind pairwise votes on pre-generated clips, with confidence intervals, cost and slice breakdowns. The design and its reasoning are in arena/DESIGN.md; the app's own notes are in arena/README.md. This Space is private and must stay private until deliberately flipped. It streams ~3,700 clips out of the private dataset multimodalart/h3-pre-gen-arena. See Going public below. hfoauth: true above creates the OAuth app and injects OAUTHCLIENTID, OAUTHCLIENTSECRET, OAUTHSCOPES and OPENIDPROVIDERURL. arena/space_auth.py implements the flow by hand (this is...

generate a video from an image with a text prompt Wan2.2 14B Fast Preview is a Hugging Face Space tagged with gradio, mcp-server, region:us. It has 1752 likes on Hugging Face.