Tools Bench.

Product launches and open-source repos with enough signal to earn a second look.

Last Brew Time: Sep 6, 2026, 10:46 AM PT

Insight

AI builders are quietly splitting into two tiers, with agent-runtime plumbing pulling ahead of everything else

Featured

GitHub242.4K

The agent that grows with you

Market Signal

Why It Has Market Pull

An extremely high-velocity open-source personal-agent project backed by Nous Research, a known AI research company. Real scale by every GitHub metric, though the sheer volume of automated-looking issue traffic makes it hard to separate genuine user feedback from bot-generated noise.

  • 242,000+ GitHub stars and nearly 50,000 forks, with new commits landing within minutes of review
  • Backed by Nous Research, the team behind the established Hermes model line, giving it organizational credibility beyond a solo project
  • Issue tracker carries tens of thousands of open items, including automated health-monitoring reports (one 'skills index degraded' thread alone has 168 comments), a sign of heavy real usage but also of noisy, partly-automated engagement
  • Real platform-compatibility complaints (Android/Termux install failures, a Codex integration regression) were filed and patched within the same release cycle

feedbacks

What People Are Saying

  • "[skills-index-watchdog] Skills index is stale or degraded (degraded)"GitHub issue

  • "fix(compression): compaction gates follow the provider's real usage, never the bytes/4 estimate"GitHub issue

  • "Add scoped Cua daemon endpoints and gateway toolset allowlists"GitHub issue

  • "[Feature]: WebUI multi-tenant login, per-user dashboard access with isolated sessions/memory"GitHub issue

  • "By v0.14.9 everyone could use Codex again."third-party review roundup

  • "a few users trying to install on phones and hitting issues with the installer script or missing binaries"third-party review roundup

GitHub205.1K

The open source coding agent.

Market Signal

Why It Has Market Pull

opencode is a large-scale, actively developed open-source coding agent backed by Anomaly (the company formerly known as SST), with adoption and revenue figures that put it in the top tier of open developer tools competing directly with commercial coding agents.

  • 205,158 GitHub stars and 26,760 forks
  • Reported over $25M ARR with subscription plans layered on top of the free tool
  • 900+ contributors and 13,000+ commits per the project's own published metrics
  • 5,746 open GitHub issues, reflecting a very large and active real-world user base
  • Backed by Anomaly Innovations, the rebranded company behind SST, OpenNext, OpenAuth, and OpenTUI

feedbacks

What People Are Saying

  • "Last update has broken the cli"GitHub issue

  • "Both the desktop and CLI versions crash frequently."GitHub issue

  • "Cli display has issues in the macOS built-in terminal"GitHub issue

  • "When Provider.resolveSDK throws errors, the session pipeline discards err.cause and stack information, making CI failures undebuggable."GitHub issue

  • "It is very similar to Claude Code in terms of capability, but offers the flexibility of being provider-agnostic."Dev.to article

  • "cli not rendering"GitHub issue

GitHub3.6K

Open source inference server that runs the best local models for your hardware, plugged into the agent you already use. Works with Pi, OpenCode, Hermes, OpenClaw, Codex, Claude Code, Oh My Pi, and Cline.

Market Signal

Why It Has Market Pull

A YC-backed developer tool with real, sustained builder momentum: an open-source local-inference layer that plugs into popular coding agents. It has active daily commits, a growing GitHub footprint, and coverage as a trending open-source project, putting it well ahead of most tools in this space on credibility.

  • 3,600+ GitHub stars and 256 forks, with commits pushed the same day it was reviewed
  • Backed by Y Combinator as a 2-person San Francisco team (Magnitude, YC company directory)
  • Featured on GitHub Trending / Trendshift as an actively rising repository
  • 21 open issues being triaged live, showing real day-to-day usage rather than a dormant repo
  • Multiple independent Show HN launches over its history, each drawing sustained comment threads

feedbacks

What People Are Saying

  • "inference ~6x slower than llama.cpp on same GGUF on Apple M4"GitHub issue

  • "agent sessions break permanently after a tool call on models with strict-alternation chat templates"GitHub issue

  • "Great looking project! I've been fumbling towards a similar approach"HN comment

  • "there's a lot of noise about different browser agents, and existing ones are slow, expensive, and inconsistent"HN post

  • "matches the performance of Claude Code at 5x lower token prices"HN post

  • "The Open-Source Inference Server With 3,300+ GitHub Stars That Runs AI Agents Locally for Free"Dev.to article

Sources

GitHub

Skills for Real Engineers. Straight from my .agents directory.

Makes your AI agent think like the laziest senior dev in the room. The best code is the code you never wrote.

Agent skill that removes signs of AI-generated writing from text

The LLVM Project is a collection of modular and reusable compiler and toolchain technologies.

38 editorial diagram types for Claude Code, Codex, and Pi. Self-contained HTML + SVG. No shadows. No Mermaid slop.

Product Hunt

Open source feature flags. Every flag is a markdown file that lives with the code it controls: what it does, why it exists, and what you decided, reviewed in a PR like any other change. Your coding agent installs Dif in one command, no account, and reads the context file so it knows what's live and what was already tried. A/B tests wait in the same files until you have traffic. Send events to your own analytics, or let Dif Cloud read the results and write the decision into the file.

Reflexio makes AI agents better with every interaction. When users correct an agent, a path fails, or something works particularly well, Reflexio turns that experience into behavior the agent can reuse next time. Instead of leaving valuable lessons buried in logs, your agent continuously learns what to repeat and what to avoid — with every learning visible, testable, and reversible. Reduce task failure rate by more than 30%, while saving tokens by more than 60%.

Ponytail is a plugin that makes coding agents write the least code that works. Before adding anything, it checks whether the change is needed, already exists, or can use the stdlib or a native API. Same behavior, fewer lines to own.

HyperProbe is how backend teams debug production issues they can't reproduce locally. Instead of adding a log line and waiting on a deploy, we let Claude Code, Codex, or Cursor drop read-only probes into a running service and capture the variable state that was never recorded. From there your agent debugs like it has a local repro, closing the bug in one sitting.

Experiential is the open source, zero markup gateway for BYOK, self-hosted and 1000+ marketplace models. It learns from your traffic to cut costs, recommend better models, and train a specialized model you own.

Every entry you write today locks at the time that you configured — 8pm by default. Once it locks, that's it: no editing, no rewriting history, no quietly polishing what you actually felt in the moment. Just an honest, raw, and unfiltered record of your day. Entries sync across your devices with your own private iCloud and never with a first/third-party server. When typing is not enough, record a squope (square video) or an audio note. at8pm is free to download (up to 2 entries per day).

YC Launch

Helping American local governments steward American industrial innovation Verdant · Summer 2026 · Government Tags: GovTech, Real Estate, Construction, Housing. Website: https://www.verdantapp.com/

Hacker News

I love research and development, you may have heard of me because of PJON (Padded Jittering Operative Network). It is a network protocol I started developing in 2010, which was recently implemented in silicon by the ETH Zurich university thanks to the research of Pius Sieber. I am excited to share with you TERMy, a terminal assistant built on top of the NPC-Forge framework. Unlike everything else being built today, TERMy does not use embeddings, machine-learning or LLMs. It runs on the CPU (even... (205 points, 45 comments).

We run a SaaS that handles petabytes of data. Our SRE team experimented with using claude, openclaw, langchain, etc. within our incident response workflows. We struggled with overflowing context, lethal trifecta vectors, hallucinations, and burned a lot of frontier tokens mostly on easy work. Approval fatigue was a challenge, and we drew a hard line at relaxing permissions in production. Long story short, we built and open-sourced AURA, a Rust-based harness specifically designed for the type of... (25 points, 6 comments).

A lot of my non-developer friends still think that AI is basically just Google or something you ask to write something and then copy paste it around... if they use it at all. When I ask them why they don't use the good stuff like Claude Cowork or its competitors, they tell me a few things: 1. Well, I don't wanna pay a bunch of money for it, and the $20 plans run out real fast. 2. I'm dealing with sensitive info (student grades, therapy notes, legal work, etc) and I don't wanna give it to anyone.... (7 points, 1 comments).

Hi HN! byteface here. myjs was a recent accident I'd like to share. Several years ago I began writing DOM API's in Python as a curiosity: https://github.com/byteface/domonic/ Every now and then I do another pass rolling in more features from the MDN pages. I have lots of unit tests but found that porting real code and libraries was a useful test of the DOM compatibility. But I didn't want to pollute the core any further as had already ported d3 and jquery. So I created a new repo called domonic-... (7 points, 0 comments).

Hi everyone! I have built an AI health device (ESP8266 + 240*240 screen). It can show AI Agents' status in real time with breathing bubble. And remind user dringking water, toilet, streth, etc. support: DeepSeek Harness, opencode, OpenClaw, Claude Code, Cursor. Hardware design, firmware and plugins are all opensource. urls are here: https://github.com/lovaxi/Rubato_Device https://github.com/lovaxi/Rubato_Plugins Type-c power supply, 2.4G wifi. Online sell on Tindie: https://www.tindie.com/produc... (6 points, 2 comments).

Hi HN, I'm Bor Shev, a composer and developer. Over the past two years I've been developing ShevtoneAudio Orchestrator. The idea is simple: instead of generating a finished piece of music and replacing the composer, Orchestrator takes the composer's own MIDI and develops it into a full orchestration. It analyzes the musical material — harmony, melody, rhythm, dynamics, structure and orchestral density — and creates an arrangement across strings, brass, percussion and other sections. The importan... (6 points, 4 comments).

HF Spaces

Real trained RL policies for the Microduck robot, running fully in the browser: MuJoCo compiled to WebAssembly steps the physics, onnxruntime-web runs the policy network at 50 Hz. No server, no backend. Two locomotion variants of the same robot are included: legs (walking, the default) and rollers (the wheeled skating variant). Press M (or hold D-pad up ~1 s on a gamepad) to switch; the roller model, meshes and policies are lazy-loaded on the first switch. | Mode | Checkpoint | What it does | |--------|-----------|--------------| | Run (legs) | BESTalphawalking.onnx | Velocity-tracking locomotion (arrows / WASD to steer) | | Sit | BESTalphasitstand.onnx | Sits down on its hull, stands back u...

Video generation with a synchronized soundtrack MiniMax-H3 — unquantized, split across two Spaces Joint video and soundtrack out of a single denoising pass, at bfloat16 with no quantization anywhere. This Space is the denoising half: the 61.73 GiB transformer and the two autoencoders. The 62.14 GiB Qwen3-VL conditioner runs in qwen3vl-conditioner, which this Space calls over the gradio API for every request. The weights are the public MiniMaxAI/MiniMax-H3 diffusers checkpoint. MiniMax-H3 is 195.9 GiB in bfloat16 and a ZeroGPU Space is evicted at 150 GB of storage. An unquantized single Space is therefore impossible, which is why quantized demos of it run NVFP4 or float8 weights. Cut the Mini...

Ultra-fast local NVFP4 video + synchronized audio generation MiniMax-H3 Ultra Fast — local conditioner + pruned NVFP4 on Blackwell Joint video and synchronized sound from MiniMax-H3, rebuilt for a single 96 GB Blackwell ZeroGPU worker. | layer | optimization | |---|---| | Weights | 12.5 GB pruned NVFP4 transformer: 20.1B effective parameters instead of 33.1B/61.7 GiB BF16. | | Compute | Native CUDA 13 NVFP4 tensor-core GEMMs through comfy-kitchen; higher-precision norms, embeddings and output heads. | | Residency | Transformer, conditioner and both VAEs remain GPU-resident during generation—no layerwise CPU offload. | | Conditioner | Local 15.7 GB Qwen3-VL NVFP4-AWQ checkpoint containing onl...

Rare disease hackathon 2026 🧬 Rare Disease, Real Kid: The MVA Hackathon 2026 The genome and clinical story here belong to a real child living with Mosaic Variegated Aneuploidy (MVA), an ultra-rare genetic condition affecting fewer than 50 people worldwide. There is currently no established treatment, and care today means managing symptoms. The family has opened the case to the research community, hoping someone can find an answer. Help us understand MVA better! The MVA Hackathon is not intended to provide general medical care, diagnosis, or professional medical advice. Rare Disease, Real Kid: MVA Hackathon 2026 is a Hugging Face Space tagged with gradio, region:us. It has 103 likes on Huggin...

generate a video from an image with a text prompt Wan2.2 14B Fast Preview is a Hugging Face Space tagged with gradio, mcp-server, region:us. It has 1809 likes on Hugging Face.

1.0K likes

Protect Birds ProtectBirds is a Hugging Face Space tagged with docker, region:us. It has 1029 likes on Hugging Face.