# backlinks.subprocess.io Base URL: https://backlinks.subprocess.io A read-only HTTP API over the Common Crawl HOST-level web graph. It answers one question in both directions: which hosts link to host X (inbound), and which hosts does X link to (outbound). No API key, no auth, no rate-limit headers. Nodes are HOSTS, not registered domains: www.example.com, blog.example.com and example.com are three separate nodes. To treat a whole site as one, use ?prefix=1 (see "Subdomains" below) — it folds a host together with all of its subdomains and returns the deduplicated union. ## Endpoints That is the complete API surface. There are no other endpoints and no other query parameters. GET /api/domains/inbound/{host} hosts that link TO {host} GET /api/domains/outbound/{host} hosts that {host} links to query params (both): ?prefix=1 ?limit=N ?offset=M GET /api JSON index metadata GET /healthz "ok" GET /llms.txt this file (The path segment is still "/domains/" for backwards compatibility; the input is a host.) Input is normalised server-side: scheme and path are stripped, the value is lowercased, and a leading "www." is removed — so stripe.com, WWW.Stripe.com and https://stripe.com/pricing all resolve to the host stripe.com. ## Examples Basic inbound (exact host — the hosts that link to the bare host botbrains.io): $ curl -s https://backlinks.subprocess.io/api/domains/inbound/botbrains.io docs.botbrains.io status.botbrains.io ssotax.org That is only three hosts because this is the HOST graph and the query is exact. To fold in the subdomains (blog., docs., www., …) as one site, add ?prefix=1 — for botbrains.io that grows the inbound set from 3 to 12 (see "Subdomains"): $ curl -s https://backlinks.subprocess.io/api/domains/inbound/botbrains.io?prefix=1 geliebte.ai community.docebo.com transform-sports.com chatgpt-prompts.de christoph-gey.de www.easybill.de botbrains.io docs.botbrains.io status.botbrains.io www.botbrains.io ssotax.org botbrains-docs.notion.site Basic outbound (whole site, ?prefix=1 — what botbrains.io and its subdomains link out to; 353 hosts in total): $ curl -s "https://backlinks.subprocess.io/api/domains/outbound/botbrains.io?prefix=1" www.aiventic.ai alhena.ai www.balto.ai convin.ai www.eesel.ai ... Always send gzip for anything large — inbound sets run to hundreds of thousands of rows and compress ~8x: $ curl -s --compressed https://backlinks.subprocess.io/api/domains/inbound/stripe.com > stripe-inbound.csv Windowing with ?limit= and ?offset= (offset is the row index into the full neighbour list; limit caps rows returned after the offset is applied): $ curl -s "https://backlinks.subprocess.io/api/domains/inbound/stripe.com?limit=5" adam.ac boardbrief.alpha.ac study.bayswater.ac www.fdp.ac en.osc.ac $ curl -s "https://backlinks.subprocess.io/api/domains/inbound/stripe.com?limit=5&offset=5" sommerfestival.ac superst.ac ted.ac www.wow.ac www.0x1.academy Read the total without downloading the body (HEAD): $ curl -sI https://backlinks.subprocess.io/api/domains/inbound/stripe.com HTTP/2 200 content-type: text/csv; charset=utf-8 x-total-count: 138129 access-control-expose-headers: X-Total-Count vary: Accept-Encoding (That is the EXACT host stripe.com. With ?prefix=1 — stripe.com plus every subdomain — the total is 770888.) Or get headers and body in one request with -D -: $ curl -s -D - "https://backlinks.subprocess.io/api/domains/inbound/stripe.com?limit=3" HTTP/2 200 content-type: text/csv; charset=utf-8 x-total-count: 138129 ... adam.ac boardbrief.alpha.ac study.bayswater.ac Subdomains — fold a host together with all of its subdomains and return the deduplicated union (default is exact host only): $ curl -s "https://backlinks.subprocess.io/api/domains/inbound/stripe.com?prefix=1" See the "Subdomains" section below for exactly what prefix matches. Index metadata: $ curl -s https://backlinks.subprocess.io/api { "release": "cc-main-2026-may-jun-jul", "level": "host", "nodes": 240380987, "arcs": 3669171001, "built_at": "2026-08-03T18:22:37Z" } Health: $ curl -s https://backlinks.subprocess.io/healthz ok ## Response shape Content-Type is `text/csv; charset=utf-8`. The body is one host per line, in normal (not reversed) notation, LF-terminated, no header row, no quoting — a single-column CSV. It is CSV rather than JSON specifically so that future columns (e.g. an edge count or a first-seen date) can be appended without breaking clients that only read column 0. Parse it as CSV, not as "a list of lines", if you want to stay forward-compatible. Response headers you care about: X-Total-Count the FULL neighbour count for this host (the size of the union when ?prefix=1), independent of ?limit / ?offset Access-Control-Expose-Headers set to X-Total-Count, so browser JS can read it Vary: Accept-Encoding Content-Encoding: gzip only when the request sent Accept-Encoding: gzip ## Pagination Large results are common: a busy host can have hundreds of thousands of inbound hosts (with ?prefix=1, more). The pattern is a fixed page size plus an advancing offset, with X-Total-Count from the first response telling you how many pages remain: offset=0 &limit=50000 offset=50000 &limit=50000 offset=100000 &limit=50000 ... until offset >= X-Total-Count $ total=$(curl -sI https://backlinks.subprocess.io/api/domains/inbound/stripe.com \ | awk -F': ' 'tolower($1)=="x-total-count"{print $2+0}') $ for off in $(seq 0 50000 $((total-1))); do curl -s --compressed "https://backlinks.subprocess.io/api/domains/inbound/stripe.com?offset=$off&limit=50000" done > stripe-inbound.csv Notes: the window is a stable slice of the same underlying array, so paging is consistent for a given index release. An offset past the end returns an empty body with a 200. limit=0 returns zero rows (useful as a header-only probe if you cannot use HEAD). ## Subdomains This is the HOST graph, so www.example.com, blog.example.com and example.com are three separate nodes with three separate neighbour lists. A query is exact by default: /inbound/example.com returns only the hosts that link to the bare host example.com, NOT the ones that link to its subdomains. Add ?prefix=1 to treat a whole site as one. It matches the host itself PLUS every host under it (any depth of subdomain) and returns the deduplicated, sorted union of all their neighbours: $ curl -s "https://backlinks.subprocess.io/api/domains/inbound/example.com?prefix=1" # = inbound(example.com) ∪ inbound(www.example.com) ∪ inbound(blog.example.com) ∪ … Precisely: ?prefix=1 on host H matches H and everything ending in ".H" — so example.com matches example.com, www.example.com, a.b.example.com. It does NOT match sibling hosts that merely share a suffix string: example.com does NOT pull in notexample.com or example.com.evil.net. X-Total-Count reflects the size of the union, and ?limit / ?offset window that union. When to use which: prefix=1 for "who links to this site/brand" (you almost always want this for a company); exact for "who links to this specific host" (e.g. only the apex, or only the www host, or one subdomain). Note on www: a leading "www." in the INPUT is stripped during normalisation, so /inbound/www.example.com is treated as /inbound/example.com. Other subdomains in the input are not stripped. (www.example.com still appears as its own row in results, and is included by ?prefix=1.) ## WHAT IT FINDS An edge A -> B exists when a page Common Crawl fetched on host A referenced host B in the SERVED HTML — the bytes as delivered, before any JavaScript runs. That means it covers more than hyperlinks: - hyperlinks (ordinary backlinks) -