Tools Bench.

Product launches and open-source repos with enough signal to earn a second look.

Last Brew Time: Aug 30, 2026, 10:40 AM PT

Insight

Builders here are spending less energy chasing a bigger model and more effort wrapping the model they already have

Featured

GitHub60.4K

AI agent skill that researches any topic across Reddit, X, YouTube, HN, Polymarket, and the web - then synthesizes a grounded summary

Market Signal

Why It Has Market Pull

last30days-skill is an AI agent skill for Claude Code that researches any topic across Reddit, X, YouTube, Hacker News, Polymarket, and the open web, then synthesizes a grounded, citation-backed summary of what communities are discussing right now. It has grown into one of the most widely adopted third-party Claude Code skills, with an active plugin-marketplace listing and a healthy fork ecosystem.

  • Roughly 60,400 GitHub stars and 5,300 forks
  • Cited engagement of 17 Reddit threads (906 upvotes) and 20 X posts (3,750 likes) in r/ClaudeCode and r/ClaudeAI
  • Independent audit scored it 79/100 on security, utility, and maintenance
  • Installable directly via the Claude Code plugin marketplace
  • Forked and mirrored by multiple independent maintainers, signaling active downstream reuse

feedbacks

What People Are Saying

  • "Finds what the community is actually upvoting and sharing, and writes you a prompt that works today, not six months ago"product description

  • "17 Reddit threads (906 upvotes) + 20 X posts (3,750 likes) from r/ClaudeCode, r/ClaudeAI"skillsindex.dev

  • "Scored 79/100 on security, utility & maintenance"skillsindex.dev

  • "Employs a 'Judge Agent' logic to weight information based on engagement metrics"product review

  • "Forked and mirrored across multiple community repos"GitHub

  • "Latest release v3.3.0"GitHub releases

GitHub46.5K

GitNexus: The Zero-Server Code Intelligence Engine - GitNexus is a client-side knowledge graph creator that runs entirely in your browser. Drop in a git repository (Github, Gitlab, Azure, Local) or ZIP file, and get an interactive knowledge graph with a built in Graph RAG Agent. Perfect for code exploration

Market Signal

Why It Has Market Pull

GitNexus is a zero-server, client-side knowledge graph engine that runs entirely in the browser: drop in a repository or ZIP and get an interactive code graph with a built-in Graph RAG agent that plugs into Cursor, Claude Code, Codex, Windsurf, and other MCP-compatible coding agents. It grew explosively at launch, and a third-party production audit measured large concrete efficiency gains for agentic coding workflows.

  • 46,500+ GitHub stars, including 2,000+ stars in a single day at launch
  • Independent production audit measured 88% fewer tool calls and 74% token savings per retrieval query in a 17-agent environment
  • Integrates via MCP with Cursor, Claude Code, Codex, Windsurf, Cline, and more
  • Licensing has already caused at least one team to switch to an alternative, a sign of real adoption friction being worked through
  • Maintainer-documented scaling limits above 10,000-50,000 files

feedbacks

What People Are Saying

  • "88% fewer tool calls, 74% token savings per retrieval query, and complete elimination of raw file reads"Pebblous production audit

  • "2,000+ stars in a single day"launch coverage

  • "Works with Cursor, Claude Code, Antigravity, Codex, Windsurf, Cline, OpenCode, CodeBuddy, Qoder"GitHub README

  • "Precomputes structure at index time -- clustering, tracing, scoring -- so tools return complete context in one call"GitHub README

  • "Over 10,000 files risks heap overflow; over 50,000 files, run overnight"maintainer note

  • "LangWatch dropped GitNexus for MIT-licensed CodeGraphContext over the PolyForm Noncommercial gray zone"community assessment

GitHub38.9K

Turn any AI agent into an AI Scientist. The #1 Agent Skills library for science, used by 190,000+ scientists worldwide. 165 ready-to-use validated skills plus 100+ scientific databases covering biology, chemistry, medicine, and drug discovery. Compatible with Cursor, Claude Code, Codex, Pi, Antigravity, and the open Agent Skills standard.

Market Signal

Why It Has Market Pull

Scientific Agent Skills, from Palo Alto-based K-Dense AI, is a large, actively maintained open-source library of validated Agent Skills and scientific database connectors that turns coding agents like Claude Code, Cursor, and Codex into research assistants for biology, chemistry, medicine, and drug discovery. It has real organic GitHub traction, a funded company behind it, and a broader commercial product line, making it one of the strongest candidates in this batch for hands-on testing.

  • 38,953 GitHub stars and 3,642 forks, with a commit pushed as recently as yesterday
  • K-Dense AI has raised $2M in funding and is building a commercial companion product, K-Dense Web, plus a free desktop app called Kady
  • 165 validated Agent Skills spanning 100+ scientific databases, compatible with Cursor, Claude Code, Codex, Pi, and Antigravity
  • Company claims 190,000+ scientist users worldwide; this figure is self-reported and could not be independently verified

feedbacks

What People Are Saying

  • "Turn any AI agent into an AI Scientist"GitHub README

  • "Turn any AI agent into a scientist with our Claude Scientific Skills. Open source and covers all sciences including biology, physics, chemistry, astronomy, quantum, data science and has extensive writing and slide capabilities."X (Twitter)

  • "reading documentation, querying databases, writing analysis scripts, processing files, creating charts, and preparing reports into capabilities that an AI Agent can discover and call"company blog

  • "K-Dense Web vs Scientific Agent Skills: Why We Built Both (And Which One You Should Use)"company blog

  • "Independent third-party coverage is still thin."evidence gap

GitHub23.5K

Open Multi-Agent Interactive Classroom — Get an immersive, multi-agent learning experience in just one click

github
Live Demo

Market Signal

Why It Has Market Pull

OpenMAIC is a fast-growing open-source project from Tsinghua University that turns any topic into a multi-agent AI classroom with speaking AI teachers, whiteboard agents, and real-time quizzes. It is backed by more than two years of in-classroom validation with 700+ real students and just shipped a v1.0.0 Pro workbench, making it one of the stronger-momentum AI education tools worth a closer look.

  • 23,600+ GitHub stars and 4,460+ forks, growing fast since its March 2026 creation
  • Actively maintained, with commits and a v1.0.0 release as recently as August 27, 2026
  • Validated in real deployment with 700+ Tsinghua University students over more than two years of research
  • 216 open issues indicate an active, engaged user base filing feedback
  • Positive social pickup calling it a leading example of where AI in education is heading

feedbacks

What People Are Saying

  • "a great example of where AI in education is heading"X post

  • "tested with real students and now open sourced by the Tsinghua team"X post

  • "validated through extensive real-world deployment with over 700 students"GitHub README

  • "adds a Pro workbench alongside the classic one-click generator"GitHub release notes

  • "Get an immersive, multi-agent learning experience in just one click"GitHub repo description

  • "Independent third-party coverage is still thin."evidence gap

Hacker News220 pts

Hi HN, we built an open source model gateway. It's a single place to manage our own self hosted, frontier, and open source models in one place. It’s is rust native, built for concurrency, and implements all the config quirks across models and providers (streaming formats, tool calls, model parameters, rate limits, and different error behavior). The gateway adds under 1 ms for BYOK requests and under 2 ms when Experiential supplies the provider key. It has every major inference provider, and 1000... (220 points, 46 comments).

Market Signal

Why It Has Market Pull

An open-source, Y Combinator-backed model gateway that gives teams a single control plane across closed, open-source, local, and custom LLMs, adding under 2ms of latency while continuously distilling cheaper, comparably capable models from real agent traffic.

  • Launched on Hacker News as a Show HN post that drew 220 points and 46 comments, a strong reception for infrastructure tooling.
  • Backed by Y Combinator as Experiential Labs, founded in 2026, focused on continual learning for agents.
  • 758+ GitHub stars with same-day commits; supports OpenAI-compatible and Anthropic Messages APIs with drop-in repointing for Claude Code, Cursor, Codex, and Aider.
  • Free bring-your-own-key pass-through for major providers, adding under 1ms latency for BYOK requests.
  • HN commenters raised substantive technical questions on caching economics and telemetry defaults, which the founders answered directly in-thread.

feedbacks

What People Are Saying

  • "Open source and no markup is the right default for a gateway."Hacker News comment

  • "if you swap between a bunch of models, you may improve performance but cost would balloon out of control"Hacker News comment

  • "model switching typically occurs at task boundaries, rarely mid-conversation"Hacker News (founder reply)

  • "The gateway adds under 1 ms for BYOK requests"Hacker News thread summary

  • "telemetry is disabled by default and merely collects anonymous usage statistics"Hacker News (founder clarification)

  • "We built open OpenRouter that turns usage into a better model"Hacker News post title

  • "the title 'open OpenRouter' could be misleading about the project's relationship to the existing OpenRouter service"Hacker News comment

Sources

GitHub

🚀🤖 Crawl4AI: Open-source LLM Friendly Web Crawler & Scraper. Don't be shy, join here: https://discord.gg/jP8KfhDhyN

Agent skill for beautiful, verifiable architecture, workflow, sequence, data-flow, and lifecycle diagrams—self-contained HTML with motion and crisp export.

Fully automatic censorship removal for language models

A framework for building realtime voice AI agents 🤖🎙️📹

Product Hunt

The Pitch Deck Analyzer gives you investor-grade feedback in minutes — trained on 25,000+ real decks and the investor decisions that followed. Upload yours and get slide-by-slide fundability feedback — stage-calibrated, checking narrative cohesion and cross-checking your claims for contradictions investors will catch. Know exactly where you lose conviction before you ever hit send. Built by 1752vc, a VC firm evaluating 4,000+ startups a year.

Hy4 Preview is a 770B MoE model (49B active, 1M context) from Tencent, built for long-horizon agentic tasks. It autonomously handles coding, game dev, and complex document analysis, running its own tests and fixing bugs before delivery.

Parse is Cohere's document vision parsing model. It transforms unstructured data in enterprise images and documents into structured data that downstream AI agents and applications can use. Handles OCR, tables/diagrams/images, and visual grounding via bounding boxes, across 9 languages. Deploy via API, cloud, or fully on-prem/air-gapped.

Coding agent can create huge code diffs. Someone still has to read it. seendiff is a local diff tool for large changes. Track what you've seen, review in chunks, and pull the coding agent in to explain its own work.

A spy satellite simulator in your browser - then you realize the sources are public and the data is real. Photorealistic 3D globe with live aircraft, ships, satellites, earthquakes, traffic, and public cameras. Talk to the planet with hands-free voice control using a realtime AI agent. Fully open source under MIT license, runs in your browser and is designed to cost little to nothing for personal exploration. No place left behind!

Agent-native, cookieless site analytics. Your coding agent measures deploys, flags changes, and suggests next moves with evidence. No dashboard, no charts to read.

Hacker News

I think agent-first chat interfaces will be a primary software modality and busy dashboard/UI will go away. I’m not sure who exactly wins it, but I want my knowledge to grow/go with me. A lot of the “knowledge” ie research, analysis, reasoning will be done by agents as the primary user. Our current notes tools & tasks management systems were built for humans… I don’t care what the 17th thing on my bug backlog is. I want to conduct agents that can execute for me and do great work. What I built Oz... (92 points, 59 comments).

The goal was to bring down the cost at the context eng. level. We do it with Layout Memoization. Instead of dumping HTML into the context window, we have built a continual learning browser harness (read only for now). We have built an early prototype for you to try out, where you can: 1. Spins up a browser instance 2. Extract any structured or tabular data from anywhere on the open-web 3. And you can do all this at the cost of a vector search Would love to hear your thoughts on this. Thanks for... (10 points, 2 comments).

OP here: this project was born out of the frustration/paranoia that AI providers are throttling their models when their server load is too high. So, I set out to model and study the problem mathematically to understand what was happening, what I found was quite surprising. The idea seems natural: as the data center demand increases momentarily through the day, throttling their models (either using quantized versions, reducing the context window or lowering the tier of the model to a smaller one)... (6 points, 0 comments).

We love Instinct, but have been increasingly worried about the data footprint we are handing over to them, and what they might do with that data. So we built an oss self-hostable version. It includes a vault that can store cards, logins, and personal information, so it can execute complex tasks on your behalf, like: "Get me two tickets to the odyssey on saturday at my nearest theatre" "Find me the best golf grip trainer and order it for me" "Read my email and find opportunities to save money by... (5 points, 0 comments).

HF Spaces

Unified memory evaluation · Results expected August 12. Agent Memory Leaderboard · 记忆之巅 A unified, open, and reproducible evaluation platform for long-term memory systems and memory-enabled agents. Agent Memory Leaderboard (AML) compares research methods and commercial products under one evaluation contract. Candidate systems implement memory Add and Search; the official platform fixes Answer, Eval, datasets, models, configurations, result review, and publication. > First public release: The inaugural verified leaderboard is expected to be published on August 12, 2026. > 首期发布: 首期经核验榜单预计将于 2026 年 8 月 12 日发布。 Results are separated along two independent dimensions. Textual and coding tasks use...

MiniMax Music 3 Studio — diffusers demo Streams full songs from lyrics + a structured caption using the MiniMaxMusic3Pipeline diffusers port. The input surface is a single Suno-inspired custom gr.HTML composer (Simple ↔ Studio modes, section-tag chips, structured-caption fields per the official prompting guide) that drives Gradio events via trigger()/props.value; styling uses only theme CSS vars so it follows the Citrus theme natively. Weights: MiniMaxAI/MiniMax-Music3 AoTI kernels: diffusers-internal-dev/MiniMax-Music3-aoti (compiled on RTX Pro 6000, matching ZeroGPU hardware) Generation streams chunk by chunk with a configurable playback headroom. The 8B language-model stage runs eager on....

Real trained RL policies for the Microduck robot, running fully in the browser: MuJoCo compiled to WebAssembly steps the physics, onnxruntime-web runs the policy network at 50 Hz. No server, no backend. Two locomotion variants of the same robot are included: legs (walking, the default) and rollers (the wheeled skating variant). Press M (or hold D-pad up ~1 s on a gamepad) to switch; the roller model, meshes and policies are lazy-loaded on the first switch. | Mode | Checkpoint | What it does | |--------|-----------|--------------| | Run (legs) | BESTalphawalking.onnx | Velocity-tracking locomotion (arrows / WASD to steer) | | Sit | BESTalphasitstand.onnx | Sits down on its hull, stands back u...

Unified text-to-image and image editing model Text-to-image and image editing demo for sensenova/SenseNova-U1.5-8B-MoT, a natively unified multimodal model (18B params, bf16) built on the NEO-unify architecture. Leave the image upload empty for text-to-image generation, or upload one or more images and write an edit instruction for image editing. Advanced options exposes denoising steps, guidance scale, timestep shift, image guidance (editing) and the seed. The step count defaults to 28 rather than the model card's 50: a fixed-seed A/B found 28 keeps composition, prompt adherence and text rendering intact — losing only some micro-texture in landscape and skin, and nothing measurable when edi...

Real-time webcam video editing (ZeroGPU) JoyAI Video Edit — Live (ZeroGPU) Real-time streaming video editing from your webcam. Built on the JoyAI Video Edit DiT + streaming VAE pipeline, served on ZeroGPU. The heavy models (DiT 16.3B, streaming VAE, MiMo-VL text encoder) are loaded at module scope under ZeroGPU's CUDA-emulation layer. Each session leases a real GPU in a @spaces.GPU fork; the main process is a thin WebSocket byte-pipe (H.264 both ways, MJPEG fallback). The custom CUDA kernels (joyomni_ops: FP8 GEMM + fused norm/rope) ship as a prebuilt cp310 / torch-2.9.1 wheel in wheels/ — matching ZeroGPU's supported stack. Attention runs on plain cuDNN SDPA (fastest on this GPU class). Che...

generate a video from an image with a text prompt Wan2.2 14B Fast Preview is a Hugging Face Space tagged with gradio, mcp-server, region:us. It has 1604 likes on Hugging Face.