Methodology
Last updated: July 27, 2026
1. What SignalCrawler actually reads
By default, a scan searches every connected Discourse forum, plus Hacker News, GitHub Issues and Stack Exchange — which of them actually get checked depends on the topic, not a fixed list run every time. We do not claim to scan "the entire internet" — that claim can't be verified, and it sets the wrong expectation. What we read is exactly this list, plus any niche forum an admin has connected and made active.
2. Two research modes, one pipeline
Find Customers scores individual posts 0–100 on how close the author is to hiring or buying — an AI model reads each post and decides which side of the transaction the author is on first (buyer or seller/competitor) before scoring anything. A post advertising a service, launching a product, or listing skills for hire scores 0, regardless of how commercially active it sounds.
Market Opportunity Research groups posts across an entire scan into signals — a recurring problem raised by more than one person, not a single complaint. A signal only forms when the same underlying issue shows up more than once.
3. Confidence is separate from the score
Every scored post also carries a confidence label — high, medium, or low — that reflects how clearly the evidence supports the score, not how strong the opportunity is. A post with unambiguous wording ("looking for a developer to fix X") gets high confidence; a vague or inferred read gets low. The two numbers answer different questions on purpose: how strong, and how sure.
4. What a signal's score is made of
An opportunity's score is not one number invented by a language model. It's built from eight named components, and the model is only ever asked to judge one narrow dimension at a time — never the final score itself:
- Demand, cross-source spread, freshness, and growth over time — pure arithmetic over the actual posts behind the signal: how many, how many distinct communities, how recent, how much more than last time.
- Pain severity, buying intent, and solution feasibility — each judged by the model, but only that one dimension per post, never the combined score.
- Competition gap — how much room looks left (see below).
One part of this is honestly a heuristic, not a market study: the competition gap comes from how often existing tools and alternatives are named inside the posts we already collected, not from an independent search of the competitive landscape. The more rivals people name, the lower that component scores — a crowded space counts against an opportunity, not for it. When nobody names an alternative we leave the component out entirely rather than scoring it full marks: not being mentioned is not the same as not existing. We'd rather tell you that plainly than dress it up as more than it is.
5. We look for reasons it might not work, too
Every signal is checked for counter-evidence — mentions in the same posts that people aren't willing to pay, are solving it manually, or that the discussion is old or confined to one community. A "growing" signal that has real counter-evidence shows both, instead of only the flattering half.
6. When there isn't enough to say anything
A scan that returns too little to form a reliable signal is marked as insufficient evidence rather than dressed up as a normal result, and the credit is refunded. A handful of weak mentions is not the same finding as a pattern backed by independent sources, and we don't blur that difference to make a report look fuller than it is.
7. Marking a score right or wrong feeds back into measuring it
Every lead can be flagged Relevant or Wrong intent. That flag isn't just stored next to the lead — it becomes a labelled example we can measure the classifier against later: precision (of what we call a lead, how much really is one) and recall (of real leads, how many we actually found), on real posts scored by real usage, not a curated demo set. We don't yet publish that number on this page — the sample is still small — but the mechanism exists and grows every time someone leaves feedback, not from us deciding it looks good enough.
8. What we don't do
- We don't invent numbers, testimonials, or review counts anywhere on this site.
- We don't have a human editor rewrite or curate individual scan results.
- We don't treat one viral post as a market, or count reposts of the same message as independent signals.
Related
See the Data Sources Policy for what we collect and how long we keep it, and Content Removal if you'd like something taken out.