Vivly Integration Roadmap
Consolidated from 6 parallel deep-research passes (May 2026). Total surface mapped: ~300 unique data sources across 6 categories. This document is the executive view — sequencing recommendations, priority tiers, critical legal/architectural rules. The 6 detailed category reports live alongside this file.
TL;DR — The Build Sequence
Phase 0 (Day 1–30): Free or near-free, low legal risk, high signal. Ship Vivly with ~25 integrations across public knowledge, dev ecosystems, filings, scholarly research, hiring signals, and federated social. Total API spend: $0–$200/mo.
Phase 1 (Month 2–4): Strategic paid layer. Add ~15 paid sources at $30–$500/mo each. Total API spend: ~$1,500/mo. Covers crypto on-chain, finance market data, scraping fallbacks (Apify/Bright Data), person enrichment (PDL, Hunter), tech-stack intelligence.
Phase 2 (Month 4–9): MCP-first distribution. Ship Vivly as an MCP server. Bundle 9 first-party remote MCP servers (GitHub, Slack, Linear, Notion, Atlassian, Google Drive, Postgres, Filesystem, Fetch). This is the most important strategic move — MCP is now the distribution channel for "internal data."
Phase 3 (Month 6–12): Internal SaaS via unified-API layer. Stand up Nango (OSS, self-host) as the OAuth + delta-sync substrate. Add Merge.dev only if customers force you to (their per-account billing is the cost trap).
Phase 4 (Year 2): Enterprise data tier. Quote-priced sources unlock at Enterprise contract size. Crunchbase, ZoomInfo, BuiltWith Team, Sensor Tower, NewsAPI.ai, Feedly Enterprise. Total API spend: $5–25k/mo blended.
Phase 0 — The "Tier Zero" Build Set (free / near-free)
These have zero or trivial cost, low legal risk, high data richness, and durable APIs. Ship every one of them before paying for anything.
Knowledge & Scholarly (the unified-intelligence backbone)
- Wikipedia REST + MediaWiki Action API — entity descriptions, summaries
- Wikidata SPARQL — 110M structured entities, CC0
- OpenAlex — 250M+ scholarly works, CC0; free $1/day credit
- Crossref REST — 180M DOI records, CC0
- Unpaywall — OA fulltext lookup
- Semantic Scholar (S2) — citation graph, embeddings
- arXiv — preprint metadata (S3 bulk mirror for fulltext)
- PubMed E-utilities — biomedical literature
- ROR + DataCite + ORCID — research entity authority
- OpenStreetMap (Overpass + Nominatim) — geo (self-host for prod)
- USPTO Open Data Portal + EPO OPS — patents (EPO free 4 GB/wk)
Public Web & News (with critical legal architecture)
- GDELT 2.0/3.0 — free global event firehose, 15-min cadence
- RSS/Atom universal — table-stakes; 80% of newsletter/blog content
- Cloudflare Radar — internet traffic stats (CC BY-NC — non-commercial only)
- Common Crawl — petabyte-scale web archive (CC0)
- Wayback Machine CDX — historical snapshots (degrading for news in 2026)
Dev Ecosystem (Vivly's GitHub-mafia ICP)
- GH Archive (BigQuery) — every public GitHub event since 2011
- GitHub REST + GraphQL + Webhooks — code/issues/PRs/releases
- deps.dev — Google's package metadata + dep graph + scorecard (BigQuery + REST)
- OSV — vulnerability DB across all ecosystems
- OpenSSF Scorecard — security health checks for 1M+ projects
- Software Heritage — universal archival source code
- ecosyste.ms suite — packages/repos/commits/funding cross-ecosystem
- npm replication feed — every package publish since registry birth
- PyPI BigQuery downloads — richest install telemetry
- Crates.io, RubyGems, Maven Central, NuGet, Go proxy, Packagist — registries
- Homebrew analytics — macOS dev-tool adoption
Forums & Communities
- Hacker News (Algolia + Firebase) — irreplaceable signal-per-byte
- Stack Exchange API + Data Dump — 180+ Q&A sites, CC-BY-SA
- Lobsters (
.jsonendpoints) — small, elite tech community - Discourse (all instances) — Apple, AWS, OpenAI, HuggingFace forums share one API surface
- Dev.to (Forem) — CC-licensed dev content
- Hashnode — GraphQL, public read
- Lemmy — federated reddit-alikes
- Mastodon — federated social
- Bluesky / AT Protocol Jetstream — lowest stability risk in social
- Telegram MTProto — public channels (via Telethon/Pyrogram)
- Discord (within authorized servers via bot)
- Reddit (free tier 100 QPM + scraper fallback)
- Product Hunt v2 — launch signals
Finance & Filings
- SEC EDGAR API — all US public filings, Form D = real-time private funding (US gov, free)
- sec-api.io — pre-parsed EDGAR ($69/mo, optional)
- FRED (St. Louis Fed) — 800K+ macro/economic time series
- Companies House (UK) — UK corporate filings (Open Government Licence)
- Etherscan V2 multichain key — 50+ EVM chains, free 5/sec
- Alchemy — free 30M CU/month EVM/Solana RPC
- QuickNode — free 10M credits/month, 60+ chains
Hiring Signals (public job board APIs — unauth, free)
- Greenhouse Job Board API — per-customer, fully public
- Lever Postings API — per-customer, fully public
- Ashby Job Postings API — per-customer, fully public
- Workable Public Jobs Feed — per-customer, fully public
Podcasts (free + canonical)
- PodcastIndex.org — 4M+ podcasts, free, Podcasting 2.0 transcripts
- iTunes Search API — canonical podcast ID directory
Startup Discovery
- Y Combinator directory (yc-oss/api) — daily-refreshed JSON
Company Intelligence
- TheirStack — $59–169/mo, job-posting-derived tech stack + hiring intent (only "Tier-A" pick that's strictly paid; included here because it's cheap enough to belong in Phase 0)
Phase 1 — Strategic Paid Layer ($30–$500/mo each)
Add when Phase 0 hits limits, ~Month 2.
| Source | Cost | Why |
|---|---|---|
| Apify Actor marketplace | $49+/mo PAYG | The scraping toolkit. 6,000 actors. Use for TikTok, Quora, Indie Hackers, AngelList, LinkedIn fallback. |
| Bright Data | $499+/mo or $8/GB | Strongest legal track record (won Meta, X, LinkedIn cases). LinkedIn datasets + custom scraping. |
| People Data Labs Pro | $98/mo (350 person credits) | Best dev DX in person/company enrichment. |
| Hunter.io | $34/mo | Domain → email lookup, 5x faster than competitors. |
| CoinGecko Pro | $35/mo | 15K+ coins, 1 credit per call regardless of params. |
| Polygon.io Developer | $79/mo | US stocks + options, real-time. |
| Tiingo Power | $30/mo | US equities EOD + IEX intraday + news. |
| EODHD | $19.99–99.99/mo | International coverage (60+ exchanges). |
| Keepa | €49/mo | Amazon 6B+ products price history. |
| Listen Notes Pro | $180/mo | 3M podcasts (alt: just PodcastIndex if budget is tight). |
| AssemblyAI | $0.12/hr batch | Best STT economics for transcription pipelines (3× cheaper than Deepgram with diarization). |
| DataForSEO SERP | $0.60/1k queries | Cheapest credible SERP API post-Bing-death. |
| Brave Search API | $5/mo entry | Independent index (not Bing-derivative). |
| Crunchbase Basic | $99/mo | Funding round data without enterprise contract. |
| Wappalyzer Pro | $250/mo | Cheaper alternative to BuiltWith Team. |
| Beehiiv + Ghost APIs | Free–$49/mo | First-party newsletter ingest. |
| PostHog Cloud | Usage-based, generous free | Product analytics extraction. |
Total Phase 1 monthly spend: ~$1,200–$1,800/mo (varies with credit usage). At this level Vivly covers ~80% of the realistic public-web data surface a paying customer expects.
Phase 2 — MCP-First Distribution Strategy
This is the single most important strategic move. MCP transferred to the Agentic AI Foundation in 2025 (Anthropic donation; stewards include GitHub, Microsoft, PulseMCP). It is the de facto distribution channel for "internal data agents."
Ship Vivly as an MCP server
- List on the official registry (registry.modelcontextprotocol.io)
- Streamable HTTP transport (SSE is sunsetting June 30, 2026 at Atlassian; others following)
- Use this as the primary distribution channel — bigger than your own SDK
Bundle 9 first-party remote MCP servers
Federate these so Vivly's "unified query" actually unifies them:
- GitHub MCP (github/github-mcp-server) — most-installed MCP, baseline
- Slack MCP (official) — channels/messages/threads, despite 2026 rate-limit cuts
- Linear MCP (official) — issues/cycles/projects (best-designed PM API)
- Notion MCP (official) — pages/databases
- Atlassian MCP (Jira + Confluence) — official remote
- Google Drive MCP (reference) — file search/read
- PostgreSQL MCP (reference) — DB introspection
- Filesystem MCP (reference) — sandboxed file access
- Fetch MCP (reference) — URL → markdown grounded retrieval
Optional Phase 2 additions (when customers ask)
HubSpot, Salesforce, Stripe, Sentry, Vercel, Supabase, Neon, Cloudflare, Figma, Brave/Tavily/Exa search MCPs, Playwright browser MCP.
Phase 3 — Internal SaaS via Unified-API Layer
For customers whose stack predates MCP, use a unified-API substrate.
Recommended primary: Nango (OSS, code-first, self-host friendly)
- 700+ APIs, OAuth + delta syncs + unified models
- Best unit economics — no per-linked-account billing
- You keep control of sync logic
Alternative / supplement: Merge.dev
- Deepest normalization across HRIS, ATS, Accounting, CRM, Ticketing
- Per-linked-account billing escalates fast — model carefully
Domain-specific unified APIs (use only when needed)
- Finch — unified HRIS/payroll (220+ systems)
- Plaid — banking, with CFPB §1033 in flux (recheck 2026)
- Codat — accounting (Xero/QBO/Sage/NetSuite)
- Apideck — pass-through model, usage-based
Direct connectors when unified APIs underdeliver
- Slack (Marketplace listing required for non-trivial use)
- Microsoft Graph — Teams/Outlook/Drive/SharePoint/Calendar (CASA audit ~$5–20k for restricted scopes)
- Google Workspace — Gmail/Drive/Docs/Sheets/Calendar/Admin SDK (CASA same)
- Salesforce REST + Bulk + Pub/Sub (you consume customer's API allocation)
- HubSpot CRM
- Zendesk, Intercom, Front (support)
- Mixpanel/Amplitude/PostHog (product analytics)
- Zoom Cloud Recording + Transcripts, Otter, Fireflies (meeting intelligence)
- Calendly (booking data)
Phase 4 — Enterprise Data Tier (Quote-Priced, Year 2)
Unlocks when Vivly has Enterprise contract revenue to justify $5–25k/mo blended.
Company / Funding Intel ($15k–$100k+/yr)
- Crunchbase Enterprise — full funding/M&A
- PitchBook — VC/PE gold standard
- CB Insights — market maps + tech trends
- Tracxn — India/APAC strength
- Dealroom — EU strength
- HG Insights — IT spend (unique data)
- BuiltWith Team ($995/mo) — historical tech stack
- Sensor Tower / AppTopia — mobile app intelligence
- G2 Enterprise — B2B reviews + buyer intent
- Trustpilot Enterprise + Connect — consumer reviews
B2B Contact / Person Data ($50k+/yr)
- ZoomInfo — 300M profiles + intent
- Apollo Org tier — $119/user/mo
- Cognism — best EU GDPR posture
- 6sense (Slintel) — intent + technographics
- Coresignal — LinkedIn-derived datasets (residual legal risk post-Proxycurl)
News & Wire (heaviest legal licensing)
- NewsAPI.ai / Event Registry ($60–600+/mo) — enriched news
- Webz.io ($500–5k+/mo) — incl. dark web archive back to 2008
- NewsCatcher ($29–399/mo)
- Feedly Enterprise (~$1,600/mo) — Leo AI enrichment
- Reuters Connect ($25–100k+/yr)
- AP API (quote-only)
- Bloomberg B-PIPE ($50–200k+/yr)
- PR Newswire / BusinessWire / GlobeNewswire (free RSS for headlines; aggregator like Benzinga/RTPR for programmatic)
Web Traffic / SEO
- SimilarWeb API ($90–200k+/yr) — display-only license
- Semrush ($499+/mo) — keyword/SERP
- Ahrefs API (~$949/mo floor) — best backlink graph
- Diffbot Knowledge Graph ($299+/mo) — 10B+ entities
Finance (enterprise terminals)
- Refinitiv (LSEG) ($22–24k/seat/yr)
- FactSet ($12k+/seat/yr)
- S&P Capital IQ Pro (~$15k/seat/yr)
- Morningstar Direct (custom)
- AlphaSense (~$15k/seat/yr) — semantic search across filings + transcripts + broker research
On-Chain Crypto Intel
- Nansen API ($150–3k+/mo)
- Glassnode Business+ (mid-$thousands/mo) — API tier is gated
- Dune Analytics Plus ($349/mo) — SQL over indexed on-chain
- Arkham Intelligence (enterprise beta)
DO NOT INTEGRATE (sources that died, locked down, or carry unacceptable risk)
| Source | Reason | Date |
|---|---|---|
| Bing Web Search API | Retired — replaced by Azure AI Foundry "Grounding with Bing" (LLM-grounding tool, not REST search) | Aug 11, 2025 |
| IEX Cloud | Shut down | Aug 31, 2024 |
| Aylien standalone | Absorbed by Quantexa, indie dev funnel dead | Feb 2023 |
| TikTok APIs | Effectively closed to commercial use; only Research API with academic affiliation | 2023–ongoing |
| Spotify Web API | Locked down 2025; ML training explicitly prohibited; audio analysis removed | May 15, 2025 |
| Instagram Basic Display | Dead — only Business/Creator paths remain | Dec 4, 2024 |
| Medium API | Repo archived 2018; Zapier integration killed Jan 2024 | Ongoing |
| Goodreads API | Shut down | Dec 2020 |
| Replit public API | Deprecated, no supported path | 2023 |
| Sourcehut | Explicit anti-ML-training ToS — respect it | 2023 |
| libraries.io | Effectively abandoned — use ecosyste.ms | Ongoing |
| Open Hub (Ohloh) | Frozen since ~2018 | — |
| Proxycurl | LinkedIn lawsuit, shut down | July 4, 2026 |
| Glassdoor public API | Enterprise partners only | 2024 |
| Indeed Publisher API | Deprecated | — |
| AngelList API | Never returned after Wellfound split | 2022 |
| Letterboxd | Explicitly denies LLM/recommendation use | — |
| Sourcegraph hosted | Cloud product killed | Feb 2024 |
| Travis CI | Effectively abandoned | 2024 |
| Yahoo Finance unofficial | Violates Yahoo ToS — personal use only | Ongoing |
| Google Scholar scraping | ToS violation; use OpenAlex/S2 instead | — |
| Stocktwits API | Partner-only since 2023 | 2023 |
| Estimize | Coverage thinning, post-acquisition uncertainty | 2024 |
| Currents API, ContextualWeb | Orphaned / inconsistent maintenance | Ongoing |
Critical Legal & Architectural Rules
1. NYT v. OpenAI Architecture
NYT v. OpenAI (allowed to proceed March 2025, focused on regurgitation of memorized training data) reshapes the news/media architecture:
- Free / scraped news → store only titles, URLs, snippets, hashes. Never permanent vector store of full text.
- Licensed APIs (Reuters/AP/Bloomberg/NewsAPI.ai) → full text only in transient cache, not embedded long-term.
- Provenance on every Vivly answer: source + license + retrieval timestamp. This is the architectural answer to regurgitation claims.
2. LinkedIn-Derived Data — Elevated Risk in 2026
Proxycurl shut down July 2026 after LinkedIn sued. hiQ v. LinkedIn protects public-data viewing but not commercial resale, fake-account scraping, or behind-auth data. All LinkedIn aggregators (Coresignal, Bright Data, ContactOut) carry residual risk. If Vivly resells profile-level data, get outside counsel.
3. DPDP / GDPR — Public Availability ≠ Lawful Basis
DPDP (in force, India) and GDPR (EU) both require lawful basis, opt-out, DSAR handling, retention policy for any EU/India personal data. Even if data is publicly scrapeable, the legal stack must be intact.
- Cleanest GDPR-defensible person data: Cognism, FullContact.
- DPDP gotcha for Vivly (Indian entity): limited "publicly available" carve-out only when individuals make it public themselves. Aggregated B2B contact data fails this test in many cases.
4. Diversify Across 2+ Providers Per Category
The Bing API death is the lesson. Stack at least two providers per critical category:
- News: GDELT + NewsAPI.ai + Webz.io
- Search/SERP: DataForSEO + Brave + (Google CSE as compliance fallback)
- Person enrichment: PDL + Hunter
- Scraping: Apify + Bright Data
- Crypto RPC: Alchemy + QuickNode + Etherscan multichain
- Stocks: Polygon + Tiingo + EODHD
- Code intel: deps.dev + ecosyste.ms + Software Heritage
5. Honor robots.txt, ai.txt, llms.txt
Publishers increasingly use per-User-Agent blocks (GPTBot, ClaudeBot, CCBot). Honor them. Vivly's crawler User-Agent should be public, identifiable, and respectful.
6. Stack Exchange CC BY-SA (Attribution + ShareAlike)
SO content carries Creative Commons obligations downstream. If Vivly surfaces SO content to users, attribute it.
Cross-Cutting 2024–2026 Changes to Track
| Change | Date | Impact |
|---|---|---|
| Bing Search API retired | Aug 11, 2025 | Migrate to Brave/DataForSEO/Exa |
| IEX Cloud sunset | Aug 31, 2024 | Replaced by Polygon/Tiingo |
| Proxycurl shutdown | Jul 4, 2026 | LinkedIn fallback = Bright Data |
| Atlassian SSE → Streamable HTTP MCP | Jun 30, 2026 | Migrate MCP transport |
| USPTO ODP Beta sunset | May 29, 2026 | Use new ODP |
| Amazon PA-API v5 retired | May 15, 2026 | Migrate to Creators API |
| Amazon SP-API $1,400/yr dev fee | Jan 31, 2026 | Budget line item |
| Slack non-Marketplace rate-limit cuts | Mar 2026 | List on Marketplace |
| Atlassian points-based rate limits | Mar 2, 2026 | Re-engineer pacing |
| Tavily acquired by Nebius | Feb 2026 | Reassess Q3 2026 |
| Brave Search API free tier killed | Feb 2026 | $5/mo entry now |
| Stack Overflow + OpenAI training deal | May 2024 | Attribution obligatory |
| Reddit API price hike | Jun 2023 | Free tier OK; commercial = enterprise |
| Wikidata Query Service partitioning | 2025–2026 | Long queries time out; use dump |
| Crossref REST rate-limit headers | Dec 2025 | Read x-rate-limit-* headers |
| OpenAlex free tier capped at $1/day | Feb 2025 | Bulk snapshot for heavy use |
| NCBI PMC E-Utilities backend migration | Feb 2026 | Schema drift coming |
| News publishers blocking Wayback Machine | Late 2025 | 87% drop in news captures May–Oct 2025 |
Source Counts by Category
| Category | Sources mapped | Detail file |
|---|---|---|
| Social, community, forums | 38+ | 01-social-community.md |
| Developer ecosystems | 50+ | 02-developer-ecosystems.md |
| News, media, video, podcast | 60+ | 03-news-media.md |
| Company intel & B2B data | 38 | 04-company-intel.md |
| Financial, market, web analytics | 83 | 05-financial-markets.md |
| Knowledge, enterprise SaaS, MCP | 70+ (incl. 25+ MCP servers, 20+ integration platforms) | 06-knowledge-enterprise-mcp.md |
| Total unique sources | ~300 |
The "Negentropy" View
If you only ship 10 integrations, ship these. Each is high-leverage, durable, and compounds with the others:
- GH Archive (BigQuery) — every public GitHub event since 2011 (free)
- SEC EDGAR — all US filings + Form D = real-time private funding (free)
- GDELT 2.0/3.0 — global news/event firehose (free)
- OpenAlex — 250M scholarly works, CC0 (free)
- Wikidata SPARQL — structured knowledge graph (free, CC0)
- PodcastIndex.org — 4M podcasts, complete graph (free)
- Hacker News + Stack Exchange — irreplaceable dev signal (free)
- Greenhouse + Lever + Ashby + Workable — hiring signals (free, public, unauth)
- Etherscan multichain + Alchemy + QuickNode — on-chain (free tiers cover most use)
- MCP server federation — GitHub + Slack + Linear + Notion + Atlassian + Google Drive (this is the distribution moat, not the data)
Total cost: $0/mo. Total signal: 60% of what Vivly's "unified intelligence" claim actually requires. The next 40% is paid (Phase 1) and enterprise (Phase 4).