Tools Bench.

Product launches and open-source repos with enough signal to earn a second look.

Last Brew Time: Sep 28, 2026, 11:37 AM PT

Insight

Builders are racing straight to the plumbing underneath the agent

Featured

GitHub40.7K

Hindsight: Agent Memory That Learns

Market Signal

Why It Has Market Pull

Hindsight is a real, actively-shipped agent-memory system from Vectorize, a company with its own product site, blog, docs, and a public benchmark paper. Development cadence and third-party interest both look genuine, though independent long-term production-usage evidence beyond the vendor's own claims is still thin.

  • About 40,800 stars and 5,500 forks on a repo created roughly 11 months ago, with 179 open issues and near-daily commits.
  • Weekly-paced releases (v0.10.1 on 2026-09-21, v0.10.0 a week earlier) show sustained, non-abandoned maintenance.
  • Backed by a published benchmark paper (arXiv 2512.12818) claiming large LongMemEval/LoCoMo accuracy gains over a plain-context baseline.
  • A companion cookbook repo and active GitHub Discussions (including cross-project integration threads) point to real early adopters, not just marketing.
  • Growth is fast for an 11-month-old project; treat vendor claims of Fortune-500 usage as self-reported until independently corroborated.

feedbacks

What People Are Saying

  • "Hindsight: Agent Memory That Learns"GitHub README

  • "the most accurate agent memory system ever tested"Vectorize blog

  • "lifts overall accuracy from 39.0% to 83.6% ... on LongMemEval"arXiv paper

  • "explored integration possibilities with another project called CogniCore"GitHub Discussion

  • "Add Hindsight - state-of-the-art AI agent memory by Vectorize"GitHub issue (awesome-claude-code)

  • "vectorize-io/hindsight — GitHub trending stats & insights"Trendshift.io

Hacker News421 pts

Hello! We’re Sid, Alex, Ketan, and Milan. We’re building Whiteboard ( https://whiteboard.dev.fast/ ), an open-source desktop app where humans and agents can architect software together in a common workspace. Here’s our repo: https://github.com/devdotfast/whiteboard . We were missing the feeling of a “whiteboard session” with another dev where you leave with a deep understanding of a system, so we built this app for ourselves. Whiteboard plugs into the tools you already use - e.g. Claude Code, Co... (421 points, 141 comments).

Market Signal

Why It Has Market Pull

Whiteboard is a Y Combinator Winter 2026-backed desktop IDE where humans and AI coding agents share a visual canvas, plugging into Claude Code and Codex. It drew one of the larger Show HN threads in recent weeks and its GitHub repo grew to over 2,000 stars within about six weeks of going public - real builder momentum, not just launch-day hype.

  • Show HN post: 421 points, 142 comments
  • GitHub repo devdotfast/whiteboard: 2,065 stars, 91 forks, created 2026-08-18, still receiving commits as of 2026-09-28
  • Backed by Y Combinator's Winter 2026 batch
  • Founder actively answered technical questions throughout the thread, including on the 'IDE' naming and lack of in-app editing
  • Caveat: some commenters question whether a non-editable canvas duplicates existing diagramming approaches (Mermaid, C4) or risks becoming a stale second source of truth; independent long-term usage evidence beyond the launch window is still thin.

feedbacks

What People Are Saying

  • "The choice of the word 'IDE' seems like it might get you in trouble, given the small 'cannot edit files' detail"HN comment

  • "We didn't end up shipping edits because we're not sure what the exact shape should be."HN comment (founder)

  • "This creates an N+1 source of truth that will get out of date as projects progress."HN comment

  • "Why would I want my agent calling an MCP when this could just be a VS Code extension?"HN comment

  • "736mb on macOS - facepalm."HN comment

  • "Without deep code review, I ended up resetting weeks of work and knocking out 50K lines."HN comment

  • "Vibe-coded landing pages are an instant 'no thanks' for me."HN comment

Product Hunt326

OpenAI expands the GPT-6 family with Sol and Luna with faster, cheaper models trained like Astra. 50% lower API pricing, near-Astra factuality/coding/computer use gains, better prompt caching (90% off cached reads), and alignment improvements. Live in ChatGPT Work, Codex, and via API.

Market Signal

Why It Has Market Pull

This is OpenAI's own next-generation model family, covered at launch by every major tech outlet (TechCrunch, VentureBeat, MacRumors) and immediately live inside ChatGPT, Codex, and the API. Developer chatter on launch-day forums points to genuine cost and quality gains, not just marketing claims.

  • Cut API pricing roughly 50% versus the prior generation (Sol: $2/$10 per million tokens vs $4/$20; Luna: $0.10/$0.50 vs $0.20/$1.20)
  • Sol reportedly makes about half as many mistakes as the prior generation on the maker's own benchmarks
  • Rolled out simultaneously across ChatGPT Work, Codex, and the API for paid tiers, plus free access to the fast model via desktop app
  • Launched the same day as a competing flagship model release, drawing direct cost/quality comparisons from developers
  • Active developer-forum thread citing real usage-volume data (billions of cached tokens) to validate the new cache-read pricing

feedbacks

What People Are Saying

  • "If you wanted intelligence tasks, the mid-tier model was cheap enough and much smarter. If you wanted performance and cost-effectiveness, the fast model was significantly better value."Hacker News comment

  • "OpenAI releases new models, slashing API costs 50% or more"VentureBeat article

  • "OpenAI launches new model family, boasting lower cost and fewer mistakes"TechCrunch article

  • "Real 30-day usage numbers show roughly 6.5B cached-read tokens against 150M input and 20M output tokens"Hacker News comment

  • "The old middle tier had become awkward once the fast model's pricing improved earlier this year"Hacker News comment

  • "Cheaper tiers get flagship-level improvements"MacRumors article

Product Hunt142

Harmony is an AI-native ITSM platform for IT teams. Instead of bolting a chatbot onto a ticketing tool, we built the helpdesk around AI agents: employees ask in Slack or Teams, and 100+ production-ready agents handle password resets, app access, onboarding, device issues and more end to end, escalating to a human only when needed. Ticketing, asset inventory, SaaS management and workflows live in one workspace, so agents have context to act. Teams hit 60%+ auto-resolution in days.

Market Signal

Why It Has Market Pull

Harmony is a well-funded, founder-credible enterprise AI product: a $34M seed led by a top-tier investor, founders who previously built and sold a company to Cisco for $500M, and a public review base that is uniformly positive. It stands out as one of the more substantiated launches in this group.

  • $34M seed round led by Lightspeed Venture Partners, announced mid-2026
  • Founders previously built Epsagon, acquired by Cisco for $500M — real repeat-founder credibility
  • Reports a 70% no-touch resolution rate across handled IT requests, a concrete operational metric
  • 4.9/5 rating on Capterra across 28 reviews, with zero neutral or negative reviews recorded
  • 100+ production-ready agents live across Slack and Microsoft Teams for IT, HR, finance, procurement, and legal requests

feedbacks

What People Are Saying

  • "Harmony lands $34M seed for AI agents inside Slack and Teams"Industry newsletter article

  • "Harmony has made working on IT-related tickets, access management and asset management a breeze and saved their team a lot of time and sanity"Capterra review

  • "Customer Service: 4.9, Ease of Use: 4.8, and Value-for-Money: 4.7"Capterra ratings

  • "Harmony Raises $34 Million For AI Workplace Support"Tech news article

  • "An always-on AI expert for IT, HR, finance, procurement, and legal service requests inside Slack and Microsoft Teams"Independent API profile

  • "Teams hit 60%+ auto-resolution in days"Product listing

Hacker News12 pts

Prathmesh, CEO of MCPJam here. Users now start in ChatGPT, Claude, Cursor, and other AI clients. They reach your product through your MCP server. That means your users often aren’t in your product anymore. You can’t see what they prompted for, how the agent interpreted it, or whether your server helped them get the result they wanted. I saw this firsthand leading MCP technical strategy at Asana, including our ChatGPT and Claude launches. We were building high-stakes enterprise integrations, but... (12 points, 9 comments).

Market Signal

Why It Has Market Pull

A mature, actively developed testing and evaluation platform for MCP servers, already used in production by Asana to ship its ChatGPT app and Claude connector. Despite a modest points count on this particular discussion thread, the underlying product shows real, sustained builder and customer momentum.

  • GitHub repository has 2,223 stars and was still being pushed to the day this was checked, indicating an active, non-abandoned project
  • Asana is a named, documented customer - its engineers describe using MCPJam daily to test tool calls, debug OAuth flows, and speed up shipping its ChatGPT app and Claude connector
  • Co-founders are former Asana engineers, one of whom worked directly on Asana's own MCP server, giving the team first-hand domain credibility
  • This discussion drew 9 comments including a direct 'happy user' testimonial
  • Supports 16+ client configurations and 170+ models per the product site, indicating broad, non-trivial scope rather than a single-integration demo

feedbacks

What People Are Saying

  • "We have been a happy user of mcpjam! Excited to see the product add more capabilities!"HN comment

  • "the problem for me has been about creating stronger evals and knowing what I should be checking for."HN comment

  • "One big thing for me was being able to test tool calls with predictable inputs and getting the UI as output."MCPJam customer case study

  • "No longer have to do a 2 minute deploy for every UI change, or wait for an LLM response."MCPJam customer case study

  • "Useful for debugging OAuth flows, seeing which step breaks, if any, and retrieving client IDs and refresh tokens after-the-fact."MCPJam customer case study

Sources

GitHub

VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages.

The Office Harness for AI Agents — Spreadsheets, Docs, Slides, Canvas, Relational Tables, and PDF in one runtime.

Multi-agent harness that runs Claude Code and Codex together as one system

64.6K

An advanced guide which might benefit you a lot 🎉 . 韩先凯的人生进阶指南 人生进阶指南 离谱的人生 人生进阶 AI学习 AI指南 韩先凯的AI学习指南 英语学习指南/英语学习教程/英语学习/学英语

Product Hunt

A new roleplay experience. Build characters and entire AI societies with memory, personality and voice that recognize players, form relationships and react to everything around them. Live on Poland's largest roleplay server. Now, add it to yours.

284

AI gives you wrong answers in the same confident tone as right ones. Cuey cross-checks your AI's answer against ChatGPT, Claude, Gemini, and 30+ other models. It surfaces where they disagree and what your model missed. No copy-pasting between tabs, no repeating yourself, no extra subscriptions. Get a second opinion before you get burned by the first.

Go works inside Gmail, Slack, your docs, and your browser nudging better phrasing, drafting replies that sound like you, and prepping meetings before you think to. Add no-code agents that run on their own schedule. Free to start.

OpenScience is an open-source AI agent for scientific research. It reads papers, writes and runs code, works in notebooks and a terminal, and sends long jobs to Modal, your own servers, or a Slurm/PBS cluster. Give it a metric and Autoresearch runs experiments, logs every one, and keeps hill-climbing. Use 30+ models through one wallet, sign in with ChatGPT, or run a local model. It ranks #1 on agentic scientific research benchmarks. Free and open source.

KiwiDesk is a tiling window manager for macOS that keeps your Mac calm and organized: every window finds its place in one of seven layouts, apps open with a shortcut, and mouse users get a Space Bar they can drag windows onto. It works with your native macOS Desktops: bind a profile to each one and it loads the moment you swipe there, with its own spaces, layouts, colors and shortcuts. Everything is set in a native Settings app, no config file needed. Free to use, source on GitHub.

Clicks Communicator is a compact Android 17 phone built around a physical keyboard and a home screen for conversations. Use it as your main phone or a companion, with full 5G and apps while keeping messages, typing, and quick replies close at hand.

YC Launch

Building AI agents to operate the world's energy infrastructure and data centers Invertix · Fall 2026 · Industrials Tags: Artificial Intelligence, Solar Power, Infrastructure. Website: https://www.invertix.ai

Hacker News

Hi HN - long-time lurker (since 2012!), first time poster. Pizza Bot is a self-hosted desktop app for Mac, Windows, and Linux that runs AI agents in the background and exposes them through an email-like UI. Finished work shows up in Unread, and anything waiting on your approval shows up in Action. It's Apache 2.0-licensed, there's no signup and no telemetry, and you bring your own model provider: Anthropic, Amazon Bedrock, Google Gemini, OpenAI, OpenRouter, or a local model through Ollama. There... (61 points, 37 comments).

Hi HN, I’m Per, founder of Scrimba (YC S20). We’ve spent the last decade teaching people how to code with an HTML-based video format. We’ve now plugged an LLM into it, so that people can create explainer videos about anything. It’s called “Scrimba Explain”. To demo this technology for Hacker News, we built HN.watch. It’s like HN, but with explainer videos instead of articles. We create them on-the-fly the first time someone clicks on a link. While there are obvious visual drawbacks of using HTML... (54 points, 14 comments).

Hi HN, I built ai·rete·rag because I kept seeing teams put an LLM in charge of decisions that need to be auditable (lending, fraud, clinical triage), then bolt on "guardrails" after the fact. It runs the two in series instead: 1. A pure-Python Rete engine evaluates YAML rules against your facts. The verdict comes only from here. Same facts, same verdict, every time, with salience-based conflict resolution. 2. RAG retrieves passages from your own policy documents, and an LLM writes a plain-Englis... (44 points, 9 comments).

Hi Hacker News! Matvey, one of the authors, is here. While building enterprise agents, we ran into a problem: the more tools you connect to the AI, the higher the chance it will run out of control and leak sensitive data. Guardrails, in theory, should prevent this, but the situation is worrying: - Non-deterministic guardrails (LLM as a judge, auto modes, etc.) are vulnerable to prompt injections, or they lack knowledge of the data, making them inefficient (~10% data leaks on our benchmarks). - E... (22 points, 11 comments).

Hi HN, this is Yarik and Vlad from VOYGR - we are building the tools for agents and apps to engage with local businesses. It all started with our own pain point at VOYGR: calling businesses to verify if they are open. We are both from Google (Maps and Search) and even there, the merchants and venues don’t keep this info updated. So we built an API and started using it in-house. On July 4th, we were driving through Portland looking for a place to eat. Google Maps was saying “Holiday hours may var... (16 points, 4 comments).

Hey HN I'm the author of Jido, an Elixir Actor & Agent SDK. I was building out an Elixir Durable Actor/Agent server and hit a wall regarding communication protocols. I reviewed all the existing protocols; A2A, ACP, AHP and nothing met my needs. I wanted a language agnostic durable actor session protocol. Agent Host Protocol (AHP) from Microsoft was close, but was centered around "chat". The future I envision for agents goes far beyond just text chat. Thus, DASP was born. I'm submitting to HN bec... (16 points, 4 comments).

HF Spaces

Benchmarks and news on various repros of TypeSafe's Jev Who is rebuilding TypeSafe's Jev (System One / RLCD) in the open? This static Space opens on the Decision Index leaderboard; the News tab tracks the artifacts in one combined grid, color-coded by kind: Decoding: parallel constrained decoding on stock models (inference technique, no new weights) Diffusion: text diffusion models run in a "Jev mode" Trained: Jev-like scoring heads and fine-tunes, weights often on the Hub, promised models listed last Prior art: "this already exists" claims Explainers: architecture speculation, explainers, benchmarks and roundups Cards sort by a trending score: ♥ likes on X + 5 × GitHub stars + 8 × Hub likes...

Interactive demo for Qwen-Image-2.1 — unified text-to-image generation and image editing with native RGBA transparency support. 📑 Blog 🤗 Model Weights 💻 GitHub Qwen-Image-2.1 is a Hugging Face Space tagged with gradio, region:us. It has 285 likes on Hugging Face.

238 likes

Fast System 1 decisions with calibrated probabilities Laya is a fast System 1 decision engine: send a state and typed questions, get typed answers with probabilities and a confidence score. It never generates text, so there is nothing to parse and nothing to hallucinate. | type | question | answer | |---|---|---| | choice | which of these options? | the option, a probability per option, confidence | | score | where on this rubric? | a position along your levels, probabilities, confidence | | noul | is this true? | the probability that it is | The tabs are the patterns people use most: support triage, email and phishing, LLM guardrails, RAG passage filtering, moderation, model routing, and a...

6-step Qwen-Image-2.1, T2I + editing, vs-base comparison Viggle Turbo v0.2.1 — 6-step Qwen-Image-2.1 A DMD-distilled student of Qwen-Image-2.1 that generates and edits images in 6 sampling steps with no classifier-free guidance, against the teacher's 40 steps: about 5× faster. On most prompts it is hard to tell apart from the base model; small, dense text (8 steps narrows the gap) and complicated edits (multi-reference composition, face swaps, identity-preserving edits) can still fall short of it. The Comparison tab shows it side by side with the base model on the official Qwen examples. v0.2 (2026-09-23): much better sample diversity than v0.1 — intra-prompt diversity 0.93× the 40-step base...

Live Mic and Multilingual Live Mic sessions automatically stop after 30 seconds. Audio File accepts recordings up to 2 minutes long; trim longer recordings before uploading. The relay enforces audio-duration limits and uses a separate /api/diarization/file/stream route for files. The Multilingual Live Mic tab uses Nemotron 3.5 multilingual streaming ASR with Nemotron-3-Diarization. It shares the microphone controls, speaker activity lanes, and transcript display with Live Mic, while using the separate /api/diarization/multilingual/stream relay. Language is detected automatically by the deployed model. Start the microphone and wait for Live before speaking. The original Live Mic, Audio File,....

Video generation with a synchronized soundtrack MiniMax-H3 — unquantized, split across two Spaces Joint video and soundtrack out of a single denoising pass, at bfloat16 with no quantization anywhere. This Space is the denoising half: the 61.73 GiB transformer and the two autoencoders. The 62.14 GiB Qwen3-VL conditioner runs in qwen3vl-conditioner, which this Space calls over the gradio API for every request. The weights are the public MiniMaxAI/MiniMax-H3 diffusers checkpoint. MiniMax-H3 is 195.9 GiB in bfloat16 and a ZeroGPU Space is evicted at 150 GB of storage. An unquantized single Space is therefore impossible, which is why quantized demos of it run NVFP4 or float8 weights. Cut the Mini...