Tools Bench.

Product launches and open-source repos with enough signal to earn a second look.

Last Brew Time: Aug 3, 2026, 7:34 PM PT

Insight

Builders in this run keep converging on squeezing frontier-grade AI onto hardware you already own

Featured

GitHub12.0K

A framework for building realtime voice AI agents 🤖🎙️📹

Market Signal

Why It Has Market Pull

LiveKit is a well-funded, mature company whose open-source Agents framework powers voice AI for some of the biggest names in tech, making this one of the strongest and most durable signals evaluated in this pass.

  • $100M Series C in January 2026 at a $1B valuation, per major tech press
  • 12,075 GitHub stars and 3,462 forks on the agents repo, with commits pushed today
  • 100+ tagged releases in the framework's version history, showing sustained release cadence
  • Customers cited in company and press materials include major AI labs, automakers, and enterprise software companies, plus 911 emergency-service operators
  • Built-in test helpers are text-only by design, so audio/telephony/noise testing needs an extra layer — a real but minor limitation

feedbacks

What People Are Saying

  • "WebRTC streaming is really hard and LiveKit provides a great developer-friendly managed solution for that"Dev.to article

  • "building a real-time AI agent that can see and hear is easier than you might think"Blog post

  • "a state-of-the-art turn detector, which leverages both audio and text semantics to identify the optimal moment to respond"Documentation review

  • "one limitation is that the built-in test helpers run text-only by design"Independent review

  • "Voice AI platform hits $1B valuation with $100M funding round"Press headline

  • "seller of voice tools to major AI labs raises $100 million"Press headline

HF Spaces365 likes

Unlimited OCR is a Hugging Face Space tagged with gradio, region:us. It has 365 likes on Hugging Face.

Market Signal

Why It Has Market Pull

A real, high-momentum release from a major research org that went viral on launch and now anchors a large open-source ecosystem — one of the strongest traction signals in this set, tempered by a documented reliability flaw worth flagging to anyone considering it for high-stakes documents.

  • GitHub repo has 21.9k stars and 2.2k forks
  • Hugging Face Space has 365 likes, the highest engagement in this set
  • Official launch post drew about 19K views; an independent researcher's post drew about 16.7K views
  • Already integrated with major inference-serving frameworks, and distributed via multiple model hubs and a major cloud platform
  • 46 open GitHub issues and 26 pull requests show active, ongoing engagement

feedbacks

What People Are Saying

  • "just made every other OCR model look outdated"X reply

  • "transcribes 40+ pages in a single pass, with zero memory growth"X reply

  • "it can literally transcribe an entire book in a single pass"X reply

  • "the Space likely runs in a lighter mode... splitting images into horizontal strips yielded significantly better results"GitHub issue

  • "it fabricates plausible text when unable to read pixels, rather than signaling unreadability"GitHub issue

  • "3B total parameters and 500M activated, yet powerful enough to transcribe 40+ pages in one pass while keeping context intact"X post

Product Hunt245

Qwen3.8-Max is Qwen’s most capable model to date, a 2.4T-parameter MoE with 95B active parameters, 1M context, and multimodal agent capabilities for coding, research, cowork, and long-horizon tasks.

Market Signal

Why It Has Market Pull

Qwen3.8-Max is a legitimate, high-profile frontier model release from Alibaba — a 2.4-trillion-parameter model that made front-page tech and financial news and reportedly moved Alibaba's stock. It has real momentum and market weight behind it, though independent third-party benchmark verification is still pending, and its listing here reflects only a fraction of its actual reach.

  • 2.4 trillion total parameters, 95B active (MoE), 1M-token context window.
  • Covered by major tech and financial press, plus multiple front-page Hacker News threads on launch day.
  • Priced at roughly 40% of a leading rival model's input-token cost and 24% of its output-token cost.
  • Alibaba's stock reportedly rallied on the announcement.
  • No independent third-party benchmark has verified the claimed rankings yet — published numbers are Alibaba's own.

feedbacks

What People Are Saying

  • "competitive with current best models from OpenAI and Anthropic, Google, and Facebook"HN comment

  • "an open-weight race between Chinese labs benefits everyone"HN comment

  • "no third party... has scored it"HN comment

  • "enthusiasm for another open-weight frontier model met fatigue over unverified benchmarks"HN post

  • "community response reflected skepticism about the unverified nature of the claims"Reddit r/LocalLLaMA

  • "shipped before any benchmark, model card, or license"HN comment

Sources

GitHub

Use Claude Code, Codex and Pi for free from your terminal, app, IDE, or phone like OpenClaw (voice supported)

Kronos: A Foundation Model for the Language of Financial Markets

DeepSeek-native AI coding agent for your terminal. Engineered around prefix-cache stability — leave it running.

Product Hunt

Managed agent as a service: launch a long-horizon AI agent in one click — Claude Code, Codex, Hermes, or OpenClaw — with full history, managed recovery, and access through WhatsApp, iMessage, Telegram, Slack, web, API developers, and CLI.

Ctruh Studio is an AI-powered no-code platform that lets anyone create, customise and publish interactive 3D experiences for websites. Generate 3D assets with AI, build immersive product showcases, virtual stores, configurators and AR experiences directly in your browser.

Airtop builds, monitors, and optimizes your Google Ads campaigns from a conversation. Keyword research, campaign creation, waste audits, and performance reporting — no expertise required.

Every meeting recorder ships your audio to the cloud and bills you monthly. yapyap does neither. It records, transcribes, names your speakers, and turns talk into summaries and action items entirely on your own machine. “Lenses” reshape each recording into whatever you need: a summary, a to-do list, or a decision log. You can install more or build your own. Prefer the cloud? Optionally connect any major provider: OpenAI, Anthropic, Groq. Either way, yapyap is yours forever. No subscription.

Snapdown turns any part of your Mac screen into structured Markdown, preserving headings, lists, tables, and text instead of flattening everything into plain OCR. Capture with one shortcut, paste anywhere, and keep every screenshot local on Apple silicon.

Appllama is a design-research platform for app creators. Search 25,000+ screens from 600+ of the App Store’s top-earning iOS apps; follow complete onboarding, paywall, home and in-app flows; inspect UI elements, colors and fonts; and see revenue, downloads and ratings in context. New apps arrive weekly, captured live from the store. Start free no card required.

YC Launch

Species-specific biopesticides built from engineered insect viruses, updatable the moment pests develop resistance. Molagri · Summer 2026 · Industrials Tags: Synthetic Biology, Biotech, Sustainability, Climate, Agriculture. Website: https://molagri.com

Find the user-facing agent failures hidden behind successful traces and evals. Buildbox · Summer 2026 · B2B Tags: Artificial Intelligence, Analytics, Enterprise Software. Website: https://heybuildbox.com

Towards a General Intelligence for Chemistry Rasyn · Summer 2026 · Healthcare Tags: AI-powered Drug Discovery, Deep Learning, Biotech. Website: https://www.rasyn.ai/

Hacker News

Hi HN! We’re Theodore and Louis, founders of Armature (YC P26). We reconstruct the entire session behind the MCP tool calls you receive, including what the user asked their agent to do and what the agent thought. You wrap your MCP in 3 lines of code (our SDK is available in Typescript, Python and Go) and start seeing in your dashboard: - All sessions reconstructed: it’s like reading the real conversation the user had inside Claude or ChatGPT! - A ranking of your MCP most popular use cases, built... (38 points, 2 comments).

Scraping modern websites has become a massive headache. You basically have two choices: pay for an expensive API like Firecrawl/Browserbase, or run a fleet of headless Chrome instances that eat 1GB of RAM per page and still get blocked by Cloudflare. I built Draco to fix this. It’s a fast, single-binary web scraper written in Rust. You point it at a URL, and it spits out perfectly clean Markdown or structured JSON for LLMs. The secret sauce is that it doesn't just boot a browser for every reques... (13 points, 9 comments).

Hey HN, I am 16y/o and have been working on Sprocket for a while. It's an open-source AI agent that beats every other agent out there at both hardware and software. And here's the best part: Sprocket can (on its own) buy anything from any website when you tell it to do so. From hardware parts to SaaS subscriptions. Sprocket retrieves best-in-class context from the web for everything it does. It is therefore incredibly reliable. The agent harness's quality, performance, and UI rival that of Codex... (124 points, 13 comments).

Hi folks, I've noticed that a lot of people here seem exhausted by the amount of AI news on the front page. I shared hcker.news here a year ago and it has since gained a ton of filtering features, including a dedicated AI filter, so I figured I should tell more people about it. The filter works in three passes: 1. Known AI-related keywords and domains are filtered automatically. 2. An agent reviews the remaining articles and removes those it identifies as AI-related. 3. I make the call on uncert... (45 points, 9 comments).

The idea of Wienerdog was born out of my experience setting up my own simple but effective system of memory, self-improving skills and hooks for Claude Code and Codex. As I was teaching my friends and colleagues how to set up their own I found myself automating more and more of my system setup and finally I decided to publish it on GitHub to help others get more out of their AI usage. So what is Wienerdog? Simply put, it is just files — no daemon, no server, no telemetry: a collection of instruc... (9 points, 2 comments).

HF Spaces

Demo of the Collection of Qwen Image Edit LoRAs Qwen-Image-Edit-2511-LoRAs-Fast is a Hugging Face Space tagged with gradio, mcp-server, region:us. It has 2225 likes on Hugging Face.

Use multiple FLUX.2-Klein LoRAs Note: This space is experimental and may log image-uploads during certain periods for performance monitoring. Always comply with HF and model Terms of Service. Any stored images are automatically deleted after 7 days. FLUX.2 Klein multi-LoRA is a Hugging Face Space tagged with gradio, mcp-server, region:us. It has 308 likes on Hugging Face.

Run complete 3.96M and 9.36M text-to-waveform models live. Live text-to-waveform inference for both Inflect v2 release models: Inflect-Micro-v2: 9.36M parameters Inflect-Nano-v2: 3.96M parameters Choose the runtime that fits your device: ZeroGPU: server-side generation in this Space, with no local model download. Browser WebGPU: private, queue-free on-device inference in the same interface, with optional streaming and a WASM compatibility fallback. Every result is synthesized live from text. There is no reference audio, prerecorded fallback, or inference-time teacher model. Use the Compare tab to run the same text, speed, variation, and seed through both checkpoints. Long input is split auto...

93 likes

Codec-native video & image understanding with Mage-VL 4B Mage-VL — codec-native streaming multimodal model Demo of microsoft/Mage-VL, a 4B codec-native vision-language model (Mage-ViT encoder trained from scratch + Qwen3-4B-Instruct-2507 decoder). Instead of decoding video into uniformly sampled frames and pushing a dense grid of patch tokens through a ViT, Mage-VL follows the structure of a video codec: it keeps every anchor (I) frame patch and only the predicted (P) frame patches where the codec spends bits — the regions carrying real motion and new detail. Those surviving patches are packed into canvases, cutting visual tokens by >75%. Image — single-image Q&A. Video — video Q&A, switchab...

Live interactive world rollout from an image 🌍 ABot-World — Interactive World Rollout Upload a single starting image and steer a live navigable world in real time. Provide a first-frame image (image-to-video seed), describe the scene, and drive the world with WASD (move & turn) / IJKL (look & pan). The model autoregressively rolls out an action-conditioned world and streams decoded frames straight to your browser. Model: acvlab/ABot-World-0-5B-LF (built on Wan2.2-TI2V-5B) Code: amap-cvlab/ABot-World Project: ABot-World This Space runs a proper live backend/infrastructure (in the spirit of Overworld/waypoint-1-5): gradio.Server exposes ZeroGPU-friendly /startgame and /stopgame API endpoints p...

73 likes

Single-image Gaussian reconstruction 🌌 InfiniSplat: Implicit Gaussian Decoding for Large-Baseline Monocular View Synthesis Conditionally accepted to SIGGRAPH Asia 2026 (Journal Track) Jiawei Wang • Hao Yu • Yongzhen Hu • Xinyi Yang • Tao Ni • Xin Zhan • Junbo Chen † Xiaowei Zhou • Ruizhen Hu • Sida Peng † Equal contribution. † Corresponding authors. > [2026-07] 🎉 InfiniSplat has been conditionally accepted to SIGGRAPH Asia 2026 (Journal Track)! > [2026-07] 🎉 Inference code for RGB-only and depth-sensor-guided 3D Gaussian reconstruction is available now! InfiniSplat supports two practical modes for single-image 3D Gaussian reconstruction: | Capability | Input | Output | | --- | --- | --- | |...