Collection
Every source is collected by a registered collector with a declared schedule, rate limit, parser version and change-detection strategy. Collectors use the least expensive reliable method, in this order:
- Official API
- Official structured download
- Plain HTTP retrieval of a public page
- Deterministic parsing
- Lightweight AI extraction, only where deterministic parsing cannot work
- Stronger AI, only if uncertainty remains
- Browser automation, only when justified
- Human review
All collectors in production today use methods 1–4. No AI model produces or alters any stored observation. Web pages are fetched with an identifying user agent (EveryIntelBot), at most once per page per day, and only where the site’s robots.txt permits.
Validation and failure handling
- Every record is validated against a schema before it is stored. Records that fail are rejected and logged with the reason; they are never partially written.
- Duplicate records within a delivery are rejected. Values already stored are compared by a content hash; unchanged values are not written again.
- Values a publisher marks as unavailable are stored as gaps and never imputed.
- Network failures are retried with backoff. A collector that fails three consecutive times is flagged in the operations console.
History and revisions
Observations are append-only; the database rejects any attempt to modify or delete a stored value. When a publisher revises a value, EveryIntel stores the revision as a new record that references the one it supersedes and preserves the prior value. This allows reconstruction of what was known at any point since collection began. Values collected in an initial historical load reflect the publisher’s data as of that load, not necessarily the original first release.
Futures positioning (CFTC)
- Source: CFTC Commitments of Traders, Legacy report, futures only, via the CFTC Public Reporting Environment API.
- Net position = non-commercial long − non-commercial short, in contracts. Net as % of open interest = net ÷ total open interest.
- A report week is marked Verified when the published components reconcile: reportable + non-reportable long = open interest, reportable + non-reportable short = open interest, and non-commercial long + spreading + commercial long = reportable long.
- 3-year percentile: mid-rank of the latest net position within the trailing 156 weekly reports, inclusive. Not shown with fewer than 52 reports.
- Findings: a positioning extreme is recorded when the percentile is ≥ 95 or ≤ 5. A positioning shift is recorded when the weekly change is at least 2.5 standard deviations from the prior year of weekly changes and at least 3% of open interest.
- The latest report is used only if it is no more than 16 days old.
U.S. macro releases (BLS)
- Source: BLS Public Data API. Monthly observations only; annual averages are excluded.
- Values BLS footnotes as preliminary are stored with that note and lower confidence; later changes are stored as revisions.
- Each run re-reads the most recent five years so that revisions, including seasonal-adjustment revisions, are captured.
- Findings: a release is recorded when a new month first appears; a revision is recorded when a month within the last ~200 days changes materially.
Software pricing
- The basket is fixed and documented; products are added only by a reviewed change.
- The extractor records every USD amount advertised on the page with its stated unit (per user) and period (month or year). It does not assert which plan a price belongs to.
- A new observation is written only when the set of price points changes. When the extractor itself is upgraded, the next observation re-baselines the series and is not reported as a vendor change.
- “Lowest paid monthly price” is inferred (the smallest non-zero monthly amount) and labelled as such.
- Price-change findings always require human review before publication.
Publication
Findings are generated by documented rules and scored for importance (0–100). Rule-generated findings from public-domain official data may be published automatically when they meet configured importance and confidence thresholds; all others wait for human review. Every published finding lists the observations it is based on. See data provenance.
Limitations
- EveryIntel describes published data. It does not forecast prices or provide investment advice, and no finding should be read as a trading signal.
- Coverage is deliberately narrow while the collection framework is proven; coverage figures on this site are always live counts.