Methodology

Last updated: July 27, 2026

1. What SignalCrawler actually reads

By default, a scan searches Hacker News, GitHub Issues, Stack Exchange, and any connected Discourse forum. Reddit can be turned on for a scan, but it is not part of the default mix. We do not claim to scan "the entire internet" — that claim can't be verified, and it sets the wrong expectation. What we read is exactly this list, plus any niche forum an admin has connected and made active.

2. Two research modes, one pipeline

Find Customers scores individual posts 0–100 on how close the author is to hiring or buying — an AI model reads each post and decides which side of the transaction the author is on first (buyer or seller/competitor) before scoring anything. A post advertising a service, launching a product, or listing skills for hire scores 0, regardless of how commercially active it sounds.

Market Opportunity Research groups posts across an entire scan into signals — a recurring problem raised by more than one person, not a single complaint. A signal only forms when the same underlying issue shows up more than once.

3. Confidence is separate from the score

Every scored post also carries a confidence label — high, medium, or low — that reflects how clearly the evidence supports the score, not how strong the opportunity is. A post with unambiguous wording ("looking for a developer to fix X") gets high confidence; a vague or inferred read gets low. The two numbers answer different questions on purpose: how strong, and how sure.

4. What a signal's score is made of

An opportunity's score is not one number invented by a language model. It is built from named components — demand frequency, cross-source spread, growth over time, buying intent, pain severity, freshness, and a competition estimate — each computed from the actual posts behind the signal, not asked from the model as a single judgment call.

One part of this is honestly a heuristic, not a market study: the competition estimate comes from how often existing tools and alternatives are mentioned inside the posts we already collected, not from an independent search of the competitive landscape. We'd rather tell you that plainly than dress it up as more than it is.

5. We look for reasons it might not work, too

Every signal is checked for counter-evidence — mentions in the same posts that people aren't willing to pay, are solving it manually, or that the discussion is old or confined to one community. A "growing" signal that has real counter-evidence shows both, instead of only the flattering half.

6. When there isn't enough to say anything

A scan that returns too little to form a reliable signal is marked as insufficient evidence rather than dressed up as a normal result, and the credit is refunded. A handful of weak mentions is not the same finding as a pattern backed by independent sources, and we don't blur that difference to make a report look fuller than it is.

7. What we don't do

  • We don't invent numbers, testimonials, or review counts anywhere on this site.
  • We don't have a human editor rewrite or curate individual scan results.
  • We don't treat one viral post as a market, or count reposts of the same message as independent signals.

Related

See the Data Sources Policy for what we collect and how long we keep it, and Content Removal if you'd like something taken out.