Tools Bench.

Product launches and open-source repos with enough signal to earn a second look.

Last Brew Time: Sep 15, 2026, 11:50 AM PT

Insight

Builders are quietly giving AI agents the paperwork trail of a full-time employee

Featured

GitHub105.6K

AI agent toolkit: unified LLM API, agent loop, TUI, coding agent CLI

Market Signal

Why It Has Market Pull

Pi is an open-source AI agent toolkit -- a unified multi-provider LLM API, an extensible agent runtime, a terminal UI, and a coding-agent CLI -- that has become one of the most widely adopted developer-facing agent frameworks of the year, with a large downstream ecosystem of forks and packages built on top of it.

  • Over 105,000 GitHub stars and 13,000+ forks in roughly 13 months, with commits still shipping daily.
  • Its terminal-UI component alone reports around 5.1 million weekly package downloads, and its extension gallery lists more than 5,400 community packages.
  • Landed on Hacker News' front page repeatedly through 2026, including a launch thread that reportedly passed 600 points.
  • Spawned an active fork ecosystem -- one derivative alone has crossed 14,000 stars -- and is reportedly the base layer for at least one other widely covered agent product.
  • Community sentiment skews strongly positive on daily-driver adoption, with the sharpest criticism aimed at task-completion reliability on harder jobs rather than the core toolkit.

feedbacks

What People Are Saying

  • "I haven't met a single person who has tried pi for a few days and not made it their daily driver."HN comment

  • "It's nothing short of invigorating to have this degree of control over something so powerful."HN comment

  • "For me, it just didn't compare in quality with Claude CLI and OpenCode. It didn't finish the job."HN comment

  • "lean System Prompt + a tuned CLAUDE.md brought back a lot of the intelligence that Opus seemed to lose"HN comment

  • "Been trying this out instead of open code with Z.ai and it was dead simple to start. I've had a generally positive experience so far."HN comment

  • "I like using it but I get the feeling the author is not benchmaxxing in ci properly"HN comment

  • "Currently the very best"HN comment

GitHub94.7K

Production-grade engineering skills for AI coding agents.

Market Signal

Why It Has Market Pull

Built by Addy Osmani (longtime Chrome engineering lead) and collaborators, agent-skills packages senior-engineer workflows -- spec-first design, security hardening, accessibility, performance review -- into 25 plain-Markdown skill files that install into 70+ different coding agents via a single CLI command. It became one of the most-discussed AI-tooling releases of the year, drawing both strong adoption and a genuine, high-profile debate about whether structured skill files actually make agents more reliable.

  • 94,000+ GitHub stars and 10,000+ forks about seven months after launch, verified directly against the live repository
  • Front-page Hacker News debate with real substance on both sides, not just hype
  • One-command install (npx skills add addyosmani/agent-skills) works across 70+ agent tools including Claude Code, Cursor, and Copilot
  • Covered independently by multiple blogs and newsletters (Medium, Substack, Dev.to) within weeks of release
  • 124 open issues and continued commits days before this review show active maintenance, not a one-off drop

feedbacks

What People Are Saying

  • "I have been using spec-kit for the last few months and it has been AMAZING in practice."HN comment

  • "The slot machine can drop any hard requirement that you specifically in your AGENTS.md"HN comment

  • "I do find that just asking the same agent to do and check it's own work is not particularly reliable."HN comment

  • "people obsess over systems for guiding agents and construct elaborate rube-goldberg machines"HN comment

  • "I just removed superpowers from my own setup...it was really just slowing things down and burning more tokens than vanilla."HN comment

  • "Addy Osmani agent-skills review: 24 production skills, 72.6k stars"Dev.to

HF Spaces528 likes

Real trained RL policies for the Microduck robot, running fully in the browser: MuJoCo compiled to WebAssembly steps the physics, onnxruntime-web runs the policy network at 50 Hz. No server, no backend. Two locomotion variants of the same robot are included: legs (walking, the default) and rollers (the wheeled skating variant). Press M (or hold D-pad up ~1 s on a gamepad) to switch; the roller model, meshes and policies are lazy-loaded on the first switch. | Mode | Checkpoint | What it does | |--------|-----------|--------------| | Run (legs) | BESTalphawalking.onnx | Velocity-tracking locomotion (arrows / WASD to steer) | | Sit | BESTalphasitstand.onnx | Sits down on its hull, stands back u...

Market Signal

Why It Has Market Pull

A fully browser-playable simulator for a $399 open-source robot duck, running the exact same reinforcement-learning policies that ship on the physical hardware, entirely client-side with no install. It has drawn mainstream tech press coverage and an active builder community reproducing and extending the training pipeline, making it one of the stronger consumer-facing robotics launches evaluated here.

  • 528 likes, the highest of any listing evaluated in this pass
  • Covered by major tech press as a notable open-source robotics launch
  • Real physical product behind it: a $399 duck-shaped biped robot with camera, lidar, and dual IMUs, shipping to buyers before year-end
  • Community has already produced independent setup guides plus a dedicated reinforcement-learning training repository
  • Policies run in real time in-browser, matching the physical robot's control interface exactly

feedbacks

What People Are Saying

  • "Hugging Face is selling a cute $399 open source duck robot"Tech press (TechCrunch)

  • "Tiny open-source robot duck can walk and learn from its mistakes"Tech press (Interesting Engineering)

  • "A $399 Open-Source 25 cm Biped You Train with Reinforcement Learning"Tech press (MarkTechPost)

  • "the duck can waddle, pick things up with its beak, get back up when it falls, crouch, and even roller skate"Tech press summary (TechCrunch)

  • "A curated list of software, simulators, policies, agent tools and coverage for the Microduck robot"GitHub README (community list)

  • "MicroDuck RL simulator with custom-trained jump policy"GitHub README (community fork)

HF Spaces441 likes

Video generation with a synchronized soundtrack MiniMax-H3 — unquantized, split across two Spaces Joint video and soundtrack out of a single denoising pass, at bfloat16 with no quantization anywhere. This Space is the denoising half: the 61.73 GiB transformer and the two autoencoders. The 62.14 GiB Qwen3-VL conditioner runs in qwen3vl-conditioner, which this Space calls over the gradio API for every request. The weights are the public MiniMaxAI/MiniMax-H3 diffusers checkpoint. MiniMax-H3 is 195.9 GiB in bfloat16 and a ZeroGPU Space is evicted at 150 GB of storage. An unquantized single Space is therefore impossible, which is why quantized demos of it run NVFP4 or float8 weights. Cut the Mini...

Market Signal

Why It Has Market Pull

The official demo from the model's own publisher for a newly released video-and-audio generation model, offering a fast turbo mode that cuts generation time roughly five-fold versus the standard process while preserving quality. Backed directly by the model's own lab and picked up by major AI trade press within weeks of release, it is the clearest, most trustworthy entry point to the underlying technology.

  • 441 likes, the second-highest in this batch, on a demo published directly by the model's own organization
  • Covered by multiple major tech-press outlets and AI video platforms within about a month of release
  • Delivers roughly a five-times speedup while the latest checkpoint resolves earlier over-sharpening and small-motion artifacts
  • Generates synchronized 2K video with native stereo audio in a single pass, a capability few open models offer
  • Serves as the canonical, moderation-compliant reference that several unofficial community forks (including a duplicate 'uncensored' variant) build on top of

feedbacks

What People Are Saying

  • "An Omni-Modal Video Model That Generates 15-Second 2K Clips With Native Stereo Audio"Tech press (MarkTechPost)

  • "4 Things To Know About Minimax H3"Tech press (Forbes)

  • "Its main strength is not simply faster generation, but greater control over characters, camera movement, references, dialogue, and sound."Reddit sentiment roundup (virse.ai)

  • "the strongest checkpoint released with much better static and small-motion shots, markedly better micro-detail, and fully resolved over-sharpening issues"HF discussion

  • "generation can be slow, VRAM and RAM requirements are confusing"Reddit sentiment roundup (leadde.ai)

  • "Real-Time Video Now Faster Than Playback"Tech press (orcarouter.ai)

Hacker News89 pts

Hey HN, Toby from Nari Labs here. We've been working on making OSS speech models super-fast. Last year, we built Dia, the first OSS text-to-speech model capable of doing natural dialogue. Since then, so many more great speech models have been released to the public. But the market is still dominated by closed source models. We think that's an inference problem. Existing systems such as vLLM / SGLang are not well suited for multimodal inference. To prove this, we built an inference engine special... (89 points, 31 comments).

Market Signal

Why It Has Market Pull

Nari Labs, the team behind the well-known open-source Dia voice model, has built a specialized inference engine for high-accuracy, low-latency speech generation and transcription, topping independent third-party voice-AI benchmarks on both speed and word-error-rate against established providers.

  • Backed by Y Combinator, with a track record from its flagship open-source voice model that has drawn nearly 20,000 GitHub stars
  • Ranked first on latency and word-error-rate on the independent Coval voice-AI leaderboard, beating named competitors on cost as well
  • Sparked a large Hacker News discussion (89 points, 31 comments) with detailed technical back-and-forth from real engineers
  • Real users reported concrete strengths (fast generation) alongside specific rough edges (a mid-clip voice switch, playback artifacts), evidence of hands-on testing
  • The new engine's repository already has over 150 stars within weeks of launch, on top of sustained interest in the company's earlier open-source releases

feedbacks

What People Are Saying

  • "You definitely need independent evals by Datapoint AI or someone who can verify your claims about TTS quality."HN comment

  • "This is awesome. Thanks for pushing the audio pareto frontier forward."HN comment

  • "For some reason it switched voices half way through a 33 second clip."HN comment

  • "All TTS generations are too fast. It's almost I'm listening to a podcast on 1.25-1.5x speed."HN comment

  • "If you're going to announce a TTS model, service, or whatever, you really need demos."HN comment

  • "The horizontal moving elements of examples become stuck and unable to be scrolled once one of them is played."HN comment

  • "Voice models are not winner take all market unlike LLM APIs."HN comment

Sources

GitHub

Enhanced ChatGPT Clone: Features Agents, MCP, Skills, DeepSeek, Anthropic, AWS, OpenAI, Responses API, Azure, Groq, o1, GPT-5, Mistral, OpenRouter, Vertex AI, Gemini, Artifacts, AI model switching, message search, Code Interpreter, langchain, DALL-E-3, OpenAPI Actions, Functions, Secure Multi-User Auth, Presets, open-source for self-hosting. Active

Run frontier MoE models on hardware you already own — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦

VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages.

Fast, efficient, battle-tested at Alibaba's scale. Hybrid architecture code review tool: deterministic pipelines + LLM Agent, precise line-level comments, built-in multi-language ruleset (NPE, thread-safety, XSS, SQL injection), OpenAI & Anthropic compatible.

Source control for agents. Use multiple coding agents, track their changes and query them in one place

Product Hunt

Your book-a-demo button converts 1-2% of visitors. The rest leave. Naoma replaces the form with an AI account executive that runs the demo right then: it walks the prospect through your live product, answers questions, qualifies them and books the meeting. It remembers returning visitors, picks up where they left off, and writes every session back to your CRM. 50,000+ demos run for B2B SaaS teams. Now self-serve: upload your product and knowledge base test the agent today.

Web Search Agents are expert web crawling and research agents for your specific domain (company enrichment, regulations research, etc.). They self-learn your use case to go deeper into the sources that matter most to you, giving your AI deeper and more relevant web context. To get started, give your AI this link: https://docs.nimbleway.com/agent-onboarding.md

284

Developers use your APIs. Now agents do too. Elva discovers APIs from code, lets you choose what each audience can use, and runs your MCP servers with auth and analytics. Track agent activity and review changes as your code evolves

Slashy is an AI-native email client for managing all your inboxes. Slashy Assistant is the AI built into it: it drafts replies in your voice, organizes email, schedules meetings, prepares briefings, and tracks follow-ups. Connect your calendar, CRM, meeting notes, and favorite tools, and it starts helping within five minutes. Use it inside Slashy or through iMessage, Slack, or a phone call. Set up automations in plain English, and Slashy proactively handles recurring work before you have to ask.

Hello Inbox helps marketing teams get more emails into the inbox instead of spam. Test your campaigns before you send, see what is hurting deliverability, and get clear recommendations for what to fix. Instead of throwing deliverability scores and technical data at you, Hello Inbox turns the results into actionable steps so you can improve inbox placement without becoming an email deliverability expert.

243

Oats is an AI meeting note-taking tool that is completely open, local, and free. Designed to not get in the way of your meetings, no bots, no subscription needed when running locally with on-device LLM. Available on both macOS and Windows. Enhanced transcription, multi-language support, speaker recognition, assessment, coaching, auto-tracking follow ups and more features via ariso.ai cloud backend.

YC Launch

Hacker News

Hey HN, we’re Nischal & Naman. We’re brothers, and together we’re building an open-source platform for simulation based testing of voice agents (try it out in 5 mins - https://docs.egma.ai/docs/get-started/quickstart , 2 min demo video - https://youtu.be/wgDWEe5UAUY ) Platforms that help you do simulation testing already exist. But they all charge a heavy premium on top of inference costs. We believe if the industry truly wants to scale simulation testing of voice agents, we need to stop chargin... (16 points, 6 comments).

We’re opening the private beta for AgentDrive, a persistent file/workspace layer for AI agents. AgentDrive gives agents: - Versioned artifacts and folders - Hosted MCP access for coding agents - Direct upload sessions for larger files - Share links and controlled public publishing - Workspace-scoped authorization and auditability - Many supported file formats such as MD, HTML, JSON, images, videos, and even spreadsheets The MCP surface is intentionally narrow and reviewed rather than exposing th... (6 points, 11 comments).

Hi everyone, Been working on Otis, an open-source ai agent that gives you one minimal experience across local and hosted open-weight models, privacy-focused by design. On setup it recommends a local model based on the hardware Otis is running on, downloads it and runs it through llama.cpp for you. Ollama, LM Studio and Nvidia PAIR are supported too. Excited for everyone to try it and all feedback is welcome! (19 points, 4 comments).

Firefox Extension/Userscript and API to get Pangram scores for all articles on the hackernews frontpage. The extension allows you to hide articles with a high score. This is about detecting posts written by LLMs, not posts about AI. Feel free to use the API to build your own tooling/readers. Big thanks to https://news.ycombinator.com/user?id=salahadawi for providing the data :). (16 points, 3 comments).

I wanted to parody where AI seems to be heading. You start as a manual inference unit pressing a button for the next token, then systems take over; agents, sub-agents, approvals, swarms. It’s Universal Paperclips for the agentic AI era, and a small exploration of where predicting the next token might eventually take us. What’s the worst that could happen? (14 points, 5 comments).

Vibeworld is a persistent cyberpunk shared world that runs in your terminal, where you can meet other developers and scientists to have company at 2 AM while screaming at your LLM :) you can connect with other scientist, chat, and includes also use vocal chat if you want to use it. There is a world and a moon, and they are completely 3D rendered and animated. In the moon there is the "complaint crater": a place where you can scream an insult against your claude or codex session and everybody fro... (7 points, 3 comments).

HF Spaces

Video generation with a synchronized soundtrack MiniMax-H3 — unquantized, split across two Spaces Joint video and soundtrack out of a single denoising pass, at bfloat16 with no quantization anywhere. This Space is the denoising half: the 61.73 GiB transformer and the two autoencoders. The 62.14 GiB Qwen3-VL conditioner runs in qwen3vl-conditioner, which this Space calls over the gradio API for every request. The weights are the public MiniMaxAI/MiniMax-H3 diffusers checkpoint. MiniMax-H3 is 195.9 GiB in bfloat16 and a ZeroGPU Space is evicted at 150 GB of storage. An unquantized single Space is therefore impossible, which is why quantized demos of it run NVFP4 or float8 weights. Cut the Mini...

Blind A/B ranking of MiniMax-H3 acceleration variants Human-judged ranking of ~26 MiniMax-H3 acceleration variants over a 200-prompt corpus, from blind pairwise votes on pre-generated clips, with confidence intervals, cost and slice breakdowns. The design and its reasoning are in arena/DESIGN.md; the app's own notes are in arena/README.md. This Space is private and must stay private until deliberately flipped. It streams ~3,700 clips out of the private dataset multimodalart/h3-pre-gen-arena. See Going public below. hfoauth: true above creates the OAuth app and injects OAUTHCLIENTID, OAUTHCLIENTSECRET, OAUTHSCOPES and OPENIDPROVIDERURL. arena/space_auth.py implements the flow by hand (this is...

Simulate a fruit fly in your browser using WebGPU kernels A standalone fruit fly connectome demo: paint neurons, stimulate the network, and watch an articulated Three.js fly respond. Vanilla JavaScript, Vite, Three.js, and @huggingface/kernels. Simulated neural activity drives crafted walking, turning, and flight animations. The movements are illustrative, not validated predictions of fly behavior. Open the address printed by Vite, then click Download & start. All neural weights, fly meshes, fonts, and kernel templates are included in public/; setup loads them into the browser and caches verified weight chunks. No API key or external model service is required. Deploy the generated dist/ dire...

90 likes

Source-grounded AI notes with citations (open source). Source-grounded, bilingual, local-first AI note taker. Turn documents, web pages, recordings, videos and YouTube links into structured notes where every bullet cites the exact source chunk it came from — so you can verify, edit and reuse notes instead of trusting a black box. Find us on Product Hunt: Lynote on Product Hunt Open source (MIT): github.com/lynote-ai/lynote-notes AI note tools usually give you a summary you cannot check. If a note says "BM25 is used for retrieval", you should be able to click through to the sentence that said it. Lynote Notes makes traceability the default contract: notes are generated from your sources, and...

Put the person from a still into a driving video, in 4 steps Upload a driving video and a character still. The clip's first frame is repainted with gpt-image-2 so it shows your character in that exact pose and framing, and that repainted frame is the only conditioning the model gets besides the clip itself — no pose estimation, no segmentation, no masks. Takes 4–10 s of driving video and renders it at 4 sampling steps; anything longer is cut to 10.1 s. The model itself only renders frame counts of the form 17k+5 and never fewer than 124 (5.17 s), so shorter clips are held out to 124 frames with a frozen last frame and the render is trimmed back — what comes out is exactly as long as what wen...

An interactive guide to 3D representations. A Hitchhiker’s Guide to the 3D Ecosystem An interactive article by Suvaditya Mukherjee, Merve Noyan, Aritra Roy Gosthipaty, and Pedro Cuenca for ML practitioners learning 3D: eight deep labs, four supporting figures and eight original lamp adaptations of the Manim sequences. Read the article or open the Space. Node 22.12+ is required. The project produces static files; it needs no server, inference provider, API key, training job or paid compute. The preview uses localhost. Rolldown is pinned to 1.2.7 because its platform binaries are complete; regenerate lockfiles in a clean directory and verify Linux bindings before updating that pin. Hugging Fac...