Launch HN: Webhound (YC S23) – Research agent that builds datasets from the web

Launch HN: Webhound (YC S23) – Research agent that builds datasets from the web

We're the team behind Webhound (https://webhound.ai), an AI agent that builds datasets from the web based on natural language prompts. You describe what you're trying to find. The agent figures out how to structure the data and where to look, then searches, extracts the results, and outputs everything in a CSV you can export.

We've set up a special no-signup version for the HN community at https://hn.webhound.ai - just click "Continue as Guest" to try it without signing up.

Here's a demo: https://youtu.be/fGaRfPdK1Sk

We started building it after getting tired of doing this kind of research manually. Open 50 tabs, copy everything into a spreadsheet, realize it's inconsistent, start over. It felt like something an LLM should be able to handle.

Some examples of how people have used it in the past month:

Competitor analysis: "Create a comparison table of internal tooling platforms (Retool, Appsmith, Superblocks, UI Bakery, BudiBase, etc) with their free plan limits, pricing tiers, onboarding experience, integrations, and how they position themselves on their landing pages." (https://www.webhound.ai/dataset/c67c96a6-9d17-4c91-b9a0-ff69...)

Lead generation: "Find Shopify stores launched recently that sell skincare products. I want the store URLs, founder names, emails, Instagram handles, and product categories." (https://www.webhound.ai/dataset/b63d148a-8895-4aab-ac34-455e...)

Pricing tracking: "Track how the free and paid plans of note-taking apps have changed over the past 6 months using official sites and changelogs. List each app with a timeline of changes and the source for each." (https://www.webhound.ai/dataset/c17e6033-5d00-4e54-baf6-8dea...)

Investor mapping: "Find VCs who led or participated in pre-seed or seed rounds for browser-based devtools startups in the past year. Include the VC name, relevant partners, contact info, and portfolio links for context." (https://www.webhound.ai/dataset/1480c053-d86b-40ce-a620-37fd...)

Research collection: "Get a list of recent arXiv papers on weak supervision in NLP. For each, include the abstract, citation count, publication date, and a GitHub repo if available." (https://www.webhound.ai/dataset/e274ca26-0513-4296-85a5-2b7b...)

Hypothesis testing: "Check if user complaints about Figma's performance on large files have increased in the last 3 months. Search forums like Hacker News, Reddit, and Figma's community site and show the most relevant posts with timestamps and engagement metrics." (https://www.webhound.ai/dataset/42b2de49-acbf-4851-bbb7-080b...)

The first version of Webhound was a single agent running on Claude 4 Sonnet. It worked, but sessions routinely cost over $1100 and it would often get lost in infinite loops. We knew that wasn't sustainable, so we started building around smaller models.

That meant adding more structure. We introduced a multi-agent system to keep it reliable and accurate. There's a main agent, a set of search agents that run subtasks in parallel, a critic agent that keeps things on track, and a validator that double-checks extracted data before saving it. We also gave it a notepad for long-term memory, which helps avoid duplicates and keeps track of what it's already seen.

After switching to Gemini 2.5 Flash and layering in the agent system, we were able to cut costs by more than 30x while also improving speed and output quality.

The system runs in two phases. First is planning, where it decides the schema, how to search, what sources to use, and how to know when it's done. Then comes extraction, where it executes the plan and gathers the data.

It uses a text-based browser we built that renders pages as markdown and extracts content directly. We tried full browser use but it was slower and less reliable. Plain text still works better for this kind of task.

We also built scheduled refreshes to keep datasets up to date and an API so you can integrate the data directly into your workflows.

Right now, everything stays in the agent's context during a run. It starts to break down around 1000-5000 rows depending on the number of attributes. We're working on a better architecture for scaling past that.

We'd love feedback, especially from anyone who's tried solving this problem or built similar tools. Happy to answer anything in the thread.

Thanks! Moe

Discussion 80 comments · 112 points · mfkhalil · 2025-09-25
Open on HN
Loading the discussion…

Domain filters

Stories from these domains are hidden from every list. Subdomains match too: blocking substack.com also hides danluu.substack.com.

    New collection

    Delete this collection?

    About YAVCHN

    YAVCHN is a reader for Hacker News and Lobsters, with articles and discussions in separate windows or Classic pages.

    Created by Paul Parks and built with PUDL.

    YAVCHN source code on GitHub

    Privacy policy · Terms of use

    Help

    Keyboard

    j / k
    Move down and up the story list. The arrow keys scroll whatever has focus.
    Enter
    Read the marked story in the article reader.
    ]
    Read the next story in the same article-reader applet. Back returns to the previous story.
    p
    Pin or unpin the marked story, which keeps it in Pinned.
    n / N
    Move to the next or previous top-level comment in the window in front.
    c
    Collapse or expand that comment.
    f
    Hide or show the story list.
    Esc
    Close a menu or this help.
    Access key m
    Go to the menu bar. Most browsers take it with Alt on Windows and Linux, and Safari with Control and Option.
    ?
    Show this help.

    Windows

    Each story opens in a window holding its article above its discussion; drag the bar between them to share the room differently. A window can be moved by its title bar, resized from any edge, snapped to a half or a corner by dragging it there, maximised, or minimised to the bar at the foot of the page. Use Window > New reader window to open an empty reader, or Story > Open in new reader window to open another reader for the current article. Docked readers keep their articles when you select another story from the sidebar. Minimized readers can be restored and reused for their site. A window's Next story link reads on down the list in the same window.

    A link in a comment or an article to another Hacker News or Lobsters thread opens that thread in a window too. A link to a single HN comment opens the comment above its replies.

    While a story's window is in front, the Story and Discussion menus in the menu bar hold its commands: pinning, Next story, sorting, collapsing every thread, jumping to the first new comment. Each window also remembers where you were in its article and discussion, so a reload, or Back to a story that Next took you past, finds your place again. Closing a window forgets it.

    The whole arrangement lives in the address, so a bookmark or a shared link brings it back, and Back undoes the last change. Moving between Hacker News, Lobsters, their lists, Pinned and Find changes only the list, and leaves the windows open.

    The list

    The pin at the start of a row keeps the story in Pinned, and the cross at its end hides it. Pinned can be narrowed by words in the title, site or author, by source, and to the stories you haven't opened yet, and ordered by when you pinned them, by points or by comments; the filters are part of the address, so a filtered view can be bookmarked. Scroll past the end of the list to load more. Domain filters, in the View menu, hide every story from a site.

    Collections are named lists of stories. Story > Add to collection files the story in front into one or more of them, and the Collections feed shows them all or one at a time; the menu that chooses collections also creates, renames, and deletes them. A note is your own text on a story. Choose Add note in a story's toolbar to write one; it saves as you type. Rows with a note carry the note mark, and the Notes feed lists every noted story and searches the text of your notes.

    Browsing view

    View > Windowed and View > Classic select the browsing view and save your default in this browser. Window view reuses a reader for each feed. Classic view opens stories and applets as pages. Open as a page is a one-off action that does not change your saved default. Use the Windowed selector to return an article to a window. Direct page links always open as pages.

    Applets

    The Applets menu in the menu bar holds three tools, each a window of its own. Replies to me takes your Hacker News user name and lists the replies to your last thirty comments and stories, checking again every three minutes while it is open, and marking what is new since you last marked them read. Look up a user opens a profile on Hacker News or Lobsters, with their submissions and recent comments, as a commenter's name in any discussion does; the bar at the top of a profile looks up someone else in the same window, and Back returns to the one before. Who is hiring? filters the posts of HN's monthly hiring threads by the words you type.

    They read only what the sites publish to everyone, so none of them asks for a login, and your user name stays in this browser anonymously and is included in retained applet state when signed in.

    Find

    Find takes any link and lists every time it was submitted to Hacker News and Lobsters, so you can read each discussion of it.

    About

    YAVCHN never sees your Hacker News or Lobsters login. The discussion is fetched from each site's public API; to vote or reply, follow the link above the discussion, or the arrow beside a comment, to the source's own site. Without a YAVCHN account, your data stays in this browser. When signed in, pins, collections, notes, blocked domains, and retained reading state are stored with your account and synchronized across devices. Hidden stories and layout stay in this browser. The privacy policy has the details.

    Open source: github.com/paulmooreparks/yavchn. Built with PUDL.