Data Sources Policy
Last updated: July 27, 2026
1. Public data only
SignalCrawler collects only content that is publicly accessible without logging in — public posts, questions, issues, and discussions. We do not access private groups, do not log into any site to read gated content, and do not bypass paywalls, CAPTCHAs intended to block automated access, or any other access control.
2. Where a scan can look
Hacker News (via its public search API), GitHub Issues (via GitHub's public Search API), Stack Exchange (via its public API), and any Discourse-based forum an administrator has explicitly connected. Reddit is supported as an opt-in source per scan, not part of the default source list, and is read through Reddit's own API terms — not through unauthorized scraping.
3. Attribution
Every post we show links back to its original URL. We do not strip attribution or present a finding as if it were our own statement rather than something someone actually wrote publicly.
4. Retention
Raw collected content — the post text and metadata from a scan — is deleted automatically 48 hours after collection. It exists only long enough for you to review the results of that specific scan.
Signals — our own aggregated analysis of a recurring pattern across scans — are kept longer, since they are our output, not a copy of anyone's text. A small number of short evidence excerpts (a handful per signal) are kept alongside a signal as proof of where the pattern came from; these are short quotes with a link to the source, not the full original post.
5. What we don't do with collected data
- We don't sell or redistribute raw collected posts as a dataset.
- We don't build profiles of individuals beyond what's needed to show a signal's evidence.
- We don't use collected content to train a model we distribute elsewhere.
6. Source health
If a source becomes unavailable or starts failing during a scan, that scan proceeds with whatever sources did respond — it is never silently presented as a complete picture. The scan result names which sources were unavailable.
7. Requesting removal
If you run one of the sites we read, or you're the author of a post that appears as evidence, see Content Removal for how to have it taken out.