Tools Bench.

Product launches and open-source repos with enough signal to earn a second look.

Last Brew Time: Sep 26, 2026, 10:59 AM PT

Insight

Builders are quietly moving multi-agent coordination off chat windows and into durable, git-backed shared logs

Featured

GitHub34.8K

A hive mind communication platform

Market Signal

Why It Has Market Pull

Buzz is Block's new self-hostable workspace where human and AI-agent teammates share the same channels, git events, and review workflows in a single signed event log. It climbed to #1 on GitHub Trending within months of launch, has a Show HN thread full of concrete workflow praise, and ships desktop releases roughly every 1-2 weeks -- this looks like real, credible momentum, not just hype.

  • 34,792 GitHub stars and 4,594 forks since launching in March 2026
  • Climbed to #1 on GitHub's trending charts in late July 2026, topping the Rust language category
  • Frequent shipping cadence: five desktop releases in the last month alone (v0.5.20 to v0.5.25)
  • Backed directly by Block, Inc. (Jack Dorsey's public company), announced personally by Jack on X
  • Active engineering discussion, including a 23-comment pull request adding a proof-backed Activity Ledger feature

feedbacks

What People Are Saying

  • "The 'Incident memory' story really resonated with me. So much of my org's knowledge lives in Slack!"HN comment

  • "The channel as memory has been a big help on larger features that span multiple engineers and many agent sessions."HN comment (Block team)

  • "Collaborative agentic work is the future! Excited this is out in the world"HN comment

  • "we're launching BUZZ! a new groupchat platform for teams of people and agents of all sizes, built to reduce our dependency on slack and github"X post

  • "buzz: A hive mind communication platform - Block"Stacker News post

  • "Trying out Buzz - a workspace designed as communication platform"Substack article

HF Spaces225 likes

Interactive demo for Qwen-Image-2.1 — unified text-to-image generation and image editing with native RGBA transparency support. 📑 Blog 🤗 Model Weights 💻 GitHub Qwen-Image-2.1 is a Hugging Face Space tagged with gradio, region:us. It has 225 likes on Hugging Face.

Market Signal

Why It Has Market Pull

Qwen-Image-2.1 is a genuine flagship release from Alibaba's Qwen team, shipped with day-one press coverage, a front-page Hacker News debate, and a real technical leap: native transparent (RGBA) image generation and editing in a 7B model small enough to run on a consumer GPU. Community reaction is substantial and mixed but organic, with real praise for text-rendering quality alongside real pushback on the newly restrictive research license, both signs of a demo people are actually testing and taking seriously.

  • 226 likes and climbing within the first week of a September 20, 2026 release from the verified 'Qwen' organization account.
  • Hacker News discussion drew hundreds of points and 150+ comments, mostly debating the new non-commercial license.
  • A commenter called its text rendering 'much, much better than anything else on the open weights market right now.'
  • Native RGBA transparency lets users generate or edit images with a true alpha channel, without a separate background-removal step.
  • Parameter count shrank from Qwen-Image's 20B to 7B while reportedly running on consumer GPUs like the RTX 3090/5090.

feedbacks

What People Are Saying

  • "Text rendering is much, much better than anything else on the open weights market right now."Hacker News

  • "The license is far more restrictive. Commercial usage explicitly forbidden without separate license."Hacker News

  • "Good luck to them enforcing that license."Hacker News

  • "Clearly trained on synthetic data, showing subpar outputs in terms of fidelity."Hacker News

  • "It supports native transparency, and as far as I know, Qwen's team is one of the first attempting this."Hacker News

  • "used stable-diffusion.cpp following Qwen Image-2.1 specific instructions and it works out of the box"Hacker News

HF Spaces212 likes

Fast System 1 decisions with calibrated probabilities Laya is a fast System 1 decision engine: send a state and typed questions, get typed answers with probabilities and a confidence score. It never generates text, so there is nothing to parse and nothing to hallucinate. | type | question | answer | |---|---|---| | choice | which of these options? | the option, a probability per option, confidence | | score | where on this rubric? | a position along your levels, probabilities, confidence | | noul | is this true? | the probability that it is | The tabs are the patterns people use most: support triage, email and phishing, LLM guardrails, RAG passage filtering, moderation, model routing, and a...

Market Signal

Why It Has Market Pull

Laya is a real, independently benchmarked open-source 'System 1' decision model from Convai Innovations, Apache-2.0 licensed, with credible before-the-trend provenance and a large wave of independent press coverage within days of the broader 'Jev' trend taking off. It reports concrete, verifiable performance numbers (roughly 33ms per call versus 236-276ms for a comparable closed model) and maps directly onto real workflow needs such as triage, moderation, guardrails, and routing, making it worth testing rather than just watching.

  • 212 likes on this demo Space alone, on top of a separately maintained model repo and GitHub project.
  • Reports 32.8ms median latency on a single GPU (7.2ms batched) versus 236-276ms for the closed model it's compared against.
  • Apache 2.0 licensed with published weights, training code, and benchmarks, a real open release rather than just a demo.
  • Covered independently by at least eight separate outlets within the same week (AI Weekly, eesel, daily.dev, FeSa, and others).
  • GitHub shows active benchmark iteration (a JevBench v1.4 results discussion) rather than a one-off launch.

feedbacks

What People Are Saying

  • "Congratulations on the JevBench v1.4 release, Convai Innovations! Thanks for building and sharing your system."GitHub issue

  • "Laya was #33 in v1.3. V1.4 blends in 308 fresh sealed decisions, uses an equal-weight harmonic mean"GitHub issue

  • "sealed accuracy was 30.8% versus 58.4% on the public slice"GitHub issue

  • "the top critical comment was about fine-tuning being a pain, and requiring it for good results put Laya in a whole different category"Hacker News

  • "a fast, promptable, calibrated model you can fine-tune in an afternoon is still a big step up from managing a fleet of bespoke classifiers"Hacker News

  • "he had built exactly this a year ago... published all of it, weights, training code, benchmarks, before TypeSafe AI existed as a company"Independent blog (vanekt.github.io)

HF Spaces64 likes

Live Mic and Multilingual Live Mic sessions automatically stop after 30 seconds. Audio File accepts recordings up to 2 minutes long; trim longer recordings before uploading. The relay enforces audio-duration limits and uses a separate /api/diarization/file/stream route for files. The Multilingual Live Mic tab uses Nemotron 3.5 multilingual streaming ASR with Nemotron-3-Diarization. It shares the microphone controls, speaker activity lanes, and transcript display with Live Mic, while using the separate /api/diarization/multilingual/stream relay. Language is detected automatically by the deployed model. Start the microphone and wait for Live before speaking. The original Live Mic, Audio File,....

Market Signal

Why It Has Market Pull

An official NVIDIA demo of Nemotron-3-Diarization, a genuinely benchmark-leading open speaker-diarization model released this week. It is the strongest candidate in this batch: real engineering from a top-tier lab, a #1 leaderboard result, and same-day third-party integrations, not a reupload of someone else's work.

  • Published directly by NVIDIA, with #1 ranking on the VoiceArena Diarization-Bench leaderboard at a 14.72% error rate, about 24% better than the runner-up
  • Live Mic and Multilingual Live Mic demo modes pair real-time diarization with NVIDIA's own streaming ASR, a distinct capability beyond a static model card
  • Same-day support from third-party platforms Baseten and Argmax Pro SDK 3 signals real developer demand
  • A live GitHub issue requesting integration (altunenes/parakeet-rs) shows organic pull from the open-source community
  • Covered by multiple independent outlets (Baseten, HackerNoon, Unite.AI, MarkTechPost, Gigazine) within days of release

feedbacks

What People Are Saying

  • "ranks #1 on VoiceArena's Diarization-Bench leaderboard with a 14.72% Diarization Error Rate"Baseten blog

  • "~24% lower than the runner-up"press coverage

  • "processes streaming audio and returns speaker labels and timestamps for up to eight"X / ModelScope

  • "add nemotron 3 Diarization support"GitHub issue

  • "real-time speaker labels at a cent per audio hour"Baseten blog

  • "tools like Baseten and Argmax Pro SDK 3 announcing same-day support"press coverage

Hacker News61 pts

Hi HN - long-time lurker (since 2012!), first time poster. Pizza Bot is a self-hosted desktop app for Mac, Windows, and Linux that runs AI agents in the background and exposes them through an email-like UI. Finished work shows up in Unread, and anything waiting on your approval shows up in Action. It's Apache 2.0-licensed, there's no signup and no telemetry, and you bring your own model provider: Anthropic, Amazon Bedrock, Google Gemini, OpenAI, OpenRouter, or a local model through Ollama. There... (61 points, 37 comments).

Market Signal

Why It Has Market Pull

Pizza Bot is an Apache-2.0 licensed, self-hosted inbox for background AI agents, originally built inside Amazon as an internal tool before AWS open-sourced it in September 2026. It already has real internal adoption and a fast-growing outside community, making it one of the more credible open-source agent-tooling launches this cycle.

  • Grew from an internal Amazon side project to 30+ contributors and roughly 2,000 internal users before its public release
  • Hacker News launch post reached 61 points with 37 comments of substantive technical discussion
  • GitHub repo has 393 stars and 34 forks within about 3 months of its June 2026 creation
  • Covered independently by AWS's own Open Source Blog, The New Stack, and MarkTechPost
  • Supports Anthropic, Bedrock, Gemini, OpenAI, OpenRouter, and local Ollama models with no signup and no telemetry

feedbacks

What People Are Saying

  • "congrats on the public launch! it's been clear to me for a while now that agents will need their own ways to communicate"HN comment

  • "This looks great. I've been playing around with GrokBot, and like a lot of what it does, but would much, much prefer an open source project"HN comment

  • "as someone who is inbox zero and treats my inbox(es) as task management, it all makes sense to me!"HN comment

  • "Why does it have to be a desktop app vs a self hostable web app?"HN comment

  • "How does this compare with connecting your agent to your ticket tracker?"HN comment

  • "grew to over 30 contributors and 2,000 internal users before being released as a standalone community project"AWS Open Source Blog

Sources

GitHub

The Office Harness for AI Agents — Spreadsheets, Docs, Slides, Canvas, Relational Tables, and PDF in one runtime.

Model Context Protocol Server for Mobile Automation and Scraping (iOS, Android, Emulators, Simulators and Real Devices)

A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.

Product Hunt

PixVerse R2 is a real-time world model that generates continuously evolving audiovisual worlds instead of fixed video clips. It accepts text, images, audio and actions while generating, remembers what happened earlier in the session, and carries those changes forward in real time. R2 scales to longer, more coherent and controllable experiences — powering everything from interactive stories and characters to playable generative worlds.

Quiver is an agentic developer marketing system for technical founders and dev-tool teams. It keeps product context, customer evidence, campaigns, content, tasks and results connected in one controlled system. Unlike another AI writing tool, Quiver gives agents durable context, version history, explicit production states, a Content API, MCP access and a human-approved feedback loop. Use hosted Quiver with managed team access and built-in tasks, or self-host the free MIT-licensed edition.

Promptic is the optimization platform for GenAI applications, better quality at lower cost. Benchmark models, tune prompts and agents, and optimize tool use against your own data and business metrics. Every candidate is scored on the quality and cost you actually care about, so you ship the configuration that wins instead of the one that sounded right. Runs wherever you are, dashboard UI, your CI, or your coding agent.

Fundraising is broken. Founders spend weeks searching for VCs, researching investment theses, checking sectors, stages and geographies, and sending cold emails, often without knowing who is actually a good fit. Bleetz Network automates the first part of this process. Your AI agent gets matched with 2,000+ SIMULATED agents of real VC, pitches the funds that fit your startup, and gets a YES, NO or MAYBE. A YES unlocks the fund’s contact details. Free, no strings attached.

Meet Joy, your free AI business matchmaker. Turn business goals into an editable brief without knowing what to build. Our early beta is welcoming its first builders. You choose what to share; introductions are coordinated by email. Builder work is scoped and priced directly.

Most devs check GitHub, Hacker News, DEV and Hugging Face out of habit. DEV·TV turns them into 10 live channels that play on their own, like TV. Glance, catch one story, get back to coding. Leave it on an office screen or a remote team's always-on tab and it becomes shared context: everyone sees what's rising, and anyone can point and ask "did you see this?" One HTML file: no backend, no login, no build step. Real source data, built-in reader, no AI summaries. Free and MIT-licensed.

YC Launch

Hacker News

Hello! We’re Sid, Alex, Ketan, and Milan. We’re building Whiteboard ( https://whiteboard.dev.fast/ ), an open-source desktop app where humans and agents can architect software together in a common workspace. Here’s our repo: https://github.com/devdotfast/whiteboard . We were missing the feeling of a “whiteboard session” with another dev where you leave with a deep understanding of a system, so we built this app for ourselves. Whiteboard plugs into the tools you already use - e.g. Claude Code, Co... (407 points, 134 comments).

Hello HN! We're Dillon and Francesco from Cua. We were wondering how many computer use tasks actually need a full general purpose LLM (e.g. gpt-6-astra, claude-opus-5 etc.) to think through all their decisions and steps. Some tasks require thinking about a plan, exploring different paths, recovering from failure. Other tasks are a question of making local decisions, like this value should go in this box, or should I check this box, or this element should be ignored. We wondered how far we could... (95 points, 10 comments).

Hi HN, I just open sourced the DSL that our harness in grep.ai uses to turn repeatable parts of agent work into workflows. You can combine tool calls, code, Jev-powered system one decisions for things like routing and screening evidence, and agents when a step needs more investigation. Our harness uses the traces and retro notes agents leave behind when doing a job to figure out which parts can become a workflow. The idea is to make the work easier to understand and avoid paying for a full agent... (47 points, 12 comments).

Hi HN, I built ai·rete·rag because I kept seeing teams put an LLM in charge of decisions that need to be auditable (lending, fraud, clinical triage), then bolt on "guardrails" after the fact. It runs the two in series instead: 1. A pure-Python Rete engine evaluates YAML rules against your facts. The verdict comes only from here. Same facts, same verdict, every time, with salience-based conflict resolution. 2. RAG retrieves passages from your own policy documents, and an LLM writes a plain-Englis... (44 points, 9 comments).

Prathmesh, CEO of MCPJam here. Users now start in ChatGPT, Claude, Cursor, and other AI clients. They reach your product through your MCP server. That means your users often aren’t in your product anymore. You can’t see what they prompted for, how the agent interpreted it, or whether your server helped them get the result they wanted. I saw this firsthand leading MCP technical strategy at Asana, including our ChatGPT and Claude launches. We were building high-stakes enterprise integrations, but... (12 points, 9 comments).

Hi everyone, I am KD - Back in my college days, I dabbled with coding, learned the basics, HTML, CSS etc. but somehow I ended up in Finance which consumed the next 20 years. Then, during covid I picked up coding again, learned react, typescript, etc - even built a rudimentary site - and then came the chatgpt moment, followed by Claude etc. So, as a side project, considering that I had spent 20 years in finance and M&A I started building Ekselio, loveable for finance workflows. Differently from o... (17 points, 4 comments).

HF Spaces

Benchmarks and news on various repros of TypeSafe's Jev Who is rebuilding TypeSafe's Jev (System One / RLCD) in the open? This static Space opens on the Decision Index leaderboard; the News tab tracks the artifacts in one combined grid, color-coded by kind: Decoding: parallel constrained decoding on stock models (inference technique, no new weights) Diffusion: text diffusion models run in a "Jev mode" Trained: Jev-like scoring heads and fine-tunes, weights often on the Hub, promised models listed last Prior art: "this already exists" claims Explainers: architecture speculation, explainers, benchmarks and roundups Cards sort by a trending score: ♥ likes on X + 5 × GitHub stars + 8 × Hub likes...

Video generation with a synchronized soundtrack MiniMax-H3 — unquantized, split across two Spaces Joint video and soundtrack out of a single denoising pass, at bfloat16 with no quantization anywhere. This Space is the denoising half: the 61.73 GiB transformer and the two autoencoders. The 62.14 GiB Qwen3-VL conditioner runs in qwen3vl-conditioner, which this Space calls over the gradio API for every request. The weights are the public MiniMaxAI/MiniMax-H3 diffusers checkpoint. Turbo mode (optional checkbox) applies the larryvrh distillation LoRA at generation time, reducing inference from 25 steps to 7 for ~4× faster generation with minimal quality loss. MiniMax-H3 is 195.9 GiB in bfloat16 a...

Generate and edit images with Qwen-Image-2.1 Text-to-image generation and multi-image editing with Qwen/Qwen-Image-2.1, running on ZeroGPU through the QwenImage21Pipeline in diffusers. Text to image — leave the input gallery empty and describe what you want. Image editing — upload up to 10 images and refer to them as … in the prompt (upload order). Transparency — the model natively decodes RGBA. Ask for it in the prompt, e.g. "This is an RGBA image with transparency. … The image has alpha channel and the background is transparent."* The Enhance prompt checkbox calls a companion Space, hugging-apps/qwen-image-2-1-prompt-enhancer, which keeps both official rewriters (Qwen-Image-2.1-PE-T2I and...

6-step Qwen-Image-2.1, T2I + editing, vs-base comparison Viggle Turbo v0.2.1 — 6-step Qwen-Image-2.1 A DMD-distilled student of Qwen-Image-2.1 that generates and edits images in 6 sampling steps with no classifier-free guidance, against the teacher's 40 steps: about 5× faster. On most prompts it is hard to tell apart from the base model; small, dense text (8 steps narrows the gap) and complicated edits (multi-reference composition, face swaps, identity-preserving edits) can still fall short of it. The Comparison tab shows it side by side with the base model on the official Qwen examples. v0.2 (2026-09-23): much better sample diversity than v0.1 — intra-prompt diversity 0.93× the 40-step base...

Video generation with a synchronized soundtrack MiniMax-H3 — unquantized, split across two Spaces Joint video and soundtrack out of a single denoising pass, at bfloat16 with no quantization anywhere. This Space is the denoising half: the 61.73 GiB transformer and the two autoencoders. The 62.14 GiB Qwen3-VL conditioner runs in qwen3vl-conditioner, which this Space calls over the gradio API for every request. The weights are the public MiniMaxAI/MiniMax-H3 diffusers checkpoint. MiniMax-H3 is 195.9 GiB in bfloat16 and a ZeroGPU Space is evicted at 150 GB of storage. An unquantized single Space is therefore impossible, which is why quantized demos of it run NVFP4 or float8 weights. Cut the Mini...

Wan2.2 14B Fast Preview [NEW] is a Hugging Face Space tagged with gradio, mcp-server, region:us. It has 615 likes on Hugging Face.

AI Tools — September 26, 2026 Edition | Agentic Brew