A hosted read API that accepts a LinkedIn profile URL and returns the profile page as structured JSON: identity, about, experience, education, skills, certifications, languages, projects, publications, honours, volunteering, courses, patents, test scores, images and provenance.
It is purely reverse-engineered and browserless. The service speaks LinkedIn's own internal web API (Voyager) over HTTPS with a member session cookie held server-side. There is no headless browser, no automation driver and no HTML parsing anywhere in the request path.
- Quick start
- Credentials
- API reference
- Response schema
- Approach
- Architecture
- Reliability and scale
- The console
- Configuration
- Testing
- Deployment
- Known limitations
- Engineering notes
- Responsible use
Requires uv and, for the console, Node 20+.
git clone <this-repo> && cd linkedin-profile-api
cp .env.example .env # then fill in the cookie header, see Credentials
uv sync --extra dev # creates .venv and installs dependencies
cd web && npm install && npm run build && cd .. # builds the console into app/static
uv run uvicorn app.main:app --port 8000Open http://127.0.0.1:8000 for the console, or call the API directly:
curl -s "http://127.0.0.1:8000/v1/profile?url=https://www.linkedin.com/in/williamhgates/" | jqOne profile, straight from the terminal, without starting the server:
uv run linkedin-api "https://www.linkedin.com/in/williamhgates/"
uv run linkedin-api "williamhgates" --meta-only # provenance and coverage onlyWith Docker (builds the console and the API into one image):
docker build -t linkedin-profile-api .
docker run -p 8000:8000 --env-file .env linkedin-profile-apiThe service authenticates as a real LinkedIn member. Nothing is committed:
.env is git-ignored, .env.example documents every variable, and platform
deployments take secrets out of band.
Recommended - the whole cookie header. In a signed-in browser: DevTools
(F12) → Network → reload → click the first document request → Request Headers →
copy the value of cookie. Then, as a single line:
LINKEDIN_COOKIE_HEADER=li_at=AQEDA…; JSESSIONID="ajax:12345…"; bcookie="v=2&…"; bscookie="v=1&…"; lidc="b=…"
USER_AGENT=<paste navigator.userAgent from the same browser>Minimum - the cookie pair. li_at and JSESSIONID, copied at the same time
from DevTools → Application → Cookies → https://www.linkedin.com:
LINKEDIN_LI_AT=AQEDA…
LINKEDIN_JSESSIONID=ajax:12345…Why the whole header is worth the extra step: LinkedIn reads device identity
from bcookie/bscookie and datacenter routing from lidc. An auth token
arriving without them looks like a different device using a copied token, and
LinkedIn responds by expiring the session server-side (observed live: a 302 to
the same URL plus Set-Cookie: li_at=delete me). This service detects that
signal and reports LINKEDIN_SESSION_EXPIRED with instructions rather than
looping.
Credential login (/uas/authenticate) is implemented but off by default
(ALLOW_PASSWORD_LOGIN=false): a failed attempt can invalidate a working cookie
session and can put the account into a verification challenge.
Check the session at any time:
curl -s localhost:8000/readyz | jq # liveness of the upstream session
curl -s localhost:8000/v1/session | jq # source, age, breaker and cache stateBase path /v1. Every response is JSON and carries X-Request-ID,
X-Cache (HIT/MISS) and X-Response-Time-Ms.
| Parameter | Type | Default | Notes |
|---|---|---|---|
url |
string, required | A LinkedIn profile URL, or the bare vanity name | |
refresh |
boolean | false |
Bypass the cache and refetch upstream |
Accepted inputs - all of these resolve to the same profile:
https://www.linkedin.com/in/williamhgates/
https://www.linkedin.com/in/williamhgates?trk=feed&originalSubdomain=us
https://in.linkedin.com/in/williamhgates
https://m.linkedin.com/in/williamhgates/
linkedin.com/in/williamhgates
williamhgates
https://www.linkedin.com/in/ACoAAA8BYqEB… (obfuscated member id)
curl -s "$HOST/v1/profile?url=williamhgates" -H "X-API-Key: $KEY"curl -s "$HOST/v1/profile" -H 'content-type: application/json' \
-d '{"url": "https://www.linkedin.com/in/williamhgates/", "refresh": false}'| Endpoint | Purpose |
|---|---|
GET /healthz |
Liveness: version, environment, uptime |
GET /readyz |
Readiness: verifies the LinkedIn session, 503 when it is unusable |
GET /v1/session |
Session source and age, circuit breaker state, cache hit rate. Never returns cookie values |
GET /v1/profile/raw?url= |
Unparsed upstream payloads, for debugging. Disabled unless EXPOSE_RAW_ENDPOINT=true |
GET /docs, /redoc, /openapi.json |
Generated API documentation |
Every failure uses one shape, with a stable code to branch on:
{
"success": false,
"request_id": "8f2c1d4b9a70",
"error": {
"code": "PROFILE_NOT_VISIBLE",
"message": "LinkedIn returned an empty profile projection for this member.",
"details": { "identifier": "someone" }
}
}| Code | HTTP | Meaning |
|---|---|---|
INVALID_PROFILE_URL |
422 | Not a /in/ member URL (company, school, or malformed) |
VALIDATION_ERROR |
422 | Missing or unknown request field |
PROFILE_NOT_FOUND |
404 | No such member |
PROFILE_NOT_VISIBLE |
403 | Member exists, details withheld from this account |
API_KEY_MISSING / API_KEY_INVALID |
401 / 403 | This API's own authentication |
RATE_LIMITED |
429 | This API's rate limit; Retry-After set |
LINKEDIN_RATE_LIMITED |
429 | LinkedIn throttled us (429 or its non-standard 999) |
LINKEDIN_SESSION_EXPIRED |
503 | Cookie no longer valid, or signed out upstream |
LINKEDIN_CHALLENGE_REQUIRED |
503 | Account needs human verification |
LINKEDIN_AUTH_FAILED |
503 | No usable session configured |
UPSTREAM_CIRCUIT_OPEN |
503 | Breaker paused upstream calls; retry_after_seconds included |
LINKEDIN_UNAVAILABLE |
502 | Unexpected upstream response |
Set API_KEYS=key1,key2 to require X-API-Key: <key> (or
Authorization: Bearer <key>). Left unset, the API is open, which is convenient
locally. The default rate limit is 30 requests per 60 seconds per key or client
address.
Designed for consumers rather than mirroring LinkedIn's internals:
- snake_case keys, explicit nulls instead of missing keys, and lists that are always lists.
- Dates split into parts plus a padded
isostring and a humanlabel;duration_monthsisnullwhen the month is unknown rather than inventing precision. - Images as rendition sets -
urlis the largest,renditions[]carries every size with dimensions. - Public URNs (
urn:li:organization:1441, noturn:li:fs_miniCompany:1441). - Humanised enums (
Native or bilingual, notNATIVE_OR_BILINGUAL). - Self-describing completeness - see
data.meta.
data.meta is the difference between "the API returned nothing for skills" and
"this member lists no skills". A full example lives in
docs/example-response.json; the authoritative
definition is /openapi.json.
1. Which surface to call. The official APIs are partner-gated and cannot read
arbitrary profiles. A browser was ruled out by the brief and is structurally
worse (seconds and hundreds of megabytes per profile, markup that changes
weekly). Server-rendered HTML is a truncated preview when logged out and an
Ember shell when logged in. What remains is the API the LinkedIn web app itself
calls: Voyager, at https://www.linkedin.com/voyager/api/….
2. Authentication is two cookies. li_at is the session; JSESSIONID
("ajax:…") must be echoed in a csrf-token header on every call. The service
takes them from configuration, validates them with one call to
/voyager/api/me, and persists the jar so a restart does not re-authenticate.
Repeated logins are the fastest way to get an account flagged.
3. The documented endpoint is gone. identity/profiles/{id}/profileView -
the call every open-source LinkedIn client is built on - now answers HTTP 410
Gone, and /profileContactInfo, /skills and /networkinfo answer 403. This
was established by probing with a live session, not assumed.
4. The current model is the Dash collection.
GET /voyager/api/identity/dash/profiles
?q=memberIdentity
&memberIdentity=<vanity>
&decorationId=com.linkedin.voyager.dash.deco.identity.profile.FullProfileWithEntities-101
One call returns the member record plus every section as an embedded tree
(profilePositionGroups, profileEducations, profileSkills, …), with
companies and schools inlined and images behind displayImageReference. Unlike
LinkedIn's GraphQL endpoints it needs no queryId hash, so it does not break on
their release schedule.
5. Strategy chain, not a single call. LinkedIn keeps more than one model
live, so the repository tries the decorated Dash collection, then an undecorated
variant, then GraphQL if a queryId is configured, then the retired legacy
projection - stopping at the first that yields a populated profile. Both
envelopes (embedded tree, and GraphQL's normalised data + included entity
graph) are reduced to one intermediate form, so a single set of pure mappers
produces the public schema either way.
6. Enrichment degrades. Contact card, paginated skills and network counts run
concurrently and are allowed to fail: a private contact card produces a warning
in data.meta.warnings, not a 5xx.
7. Behave like a browser, or lose the session. Requests carry the header set
the web client sends (CSRF token, x-restli-protocol-version, x-li-track, a
fresh x-li-page-instance UUID, the profile's own referer). Cookies live in a
single jar - passing them per request as well made httpx emit two JSESSIONID
values, which LinkedIn answers with CSRF check failed. Rotated cookies are
absorbed back. No HTML page is ever fetched when a cookie pair is supplied.
Full reasoning, including what was tried and rejected, is in
docs/01-reverse-engineering.md.
Dependencies point one way: HTTP → services → LinkedIn integration. Nothing below the API layer knows FastAPI exists, and the parsers are pure functions.
app/
├─ api/ HTTP surface
│ ├─ routes/ profile lookups, health, readiness, session diagnostics
│ ├─ errors.py error taxonomy and handlers (one envelope for every failure)
│ ├─ security.py API key authentication (constant-time)
│ └─ ratelimit.py fixed-window limiter
├─ services/
│ ├─ profile_service.py cache lookup, request coalescing, response metadata
│ └─ cache.py TTL + LRU cache with optional disk persistence
├─ linkedin/ the reverse-engineered integration
│ ├─ endpoints.py every upstream URL, in one auditable file
│ ├─ auth.py cookie and credential flows, header set, sign-out detection
│ ├─ session.py session value object, atomic persistence
│ ├─ client.py HTTP/2 pool, retries, re-auth, status → exception mapping
│ ├─ throttle.py pacing with jitter, concurrency ceiling, circuit breaker
│ ├─ repository.py strategy chain + concurrent enrichment
│ ├─ normalized.py resolver for the GraphQL entity-graph envelope
│ └─ parsers/ pure mappers → the published schema
│ ├─ dash.py current model (embedded and normalised envelopes)
│ ├─ profile.py retired legacy projection
│ ├─ contact.py contact card
│ ├─ draft.py shared intermediate representation
│ └─ assembler.py draft → public document, with provenance
├─ models/ response schema (Pydantic v2)
├─ config.py typed settings, the only place env vars are read
└─ static/ built console (generated)
Request lifecycle:
GET /v1/profile?url=…
├─ middleware: request id, timing, structured log
├─ API key + rate limit
├─ parse and canonicalise the URL → 422 before any network call
├─ cache lookup (canonical identifier) → HIT: return, X-Cache: HIT
├─ single-flight guard → duplicate concurrent lookups share one fetch
├─ strategy chain → dashProfile → dashProfileMinimal → graphql → profileView
├─ enrichment (concurrent, failures degrade to warnings)
├─ pure parsers → draft → public document
└─ cache store → envelope
The scarce resource is not CPU - it is the LinkedIn account's request budget.
| Mechanism | Default |
|---|---|
| Response cache (TTL + LRU, optional disk persistence) | 900 s, 512 entries |
| Single-flight coalescing (N concurrent lookups → 1 upstream fetch) | always on |
| Minimum interval + jitter between upstream calls | 1.2 s + up to 0.6 s |
| Concurrency ceiling | 4 in flight |
| Retries with exponential backoff (429/5xx only; 410 never retried) | 3 attempts |
| Single re-authentication on a rejected session | per call |
| Circuit breaker (closed → open → half-open trial) | 4 failures, 90 s cooldown |
| Session persistence across restarts | on |
| This API's own rate limit | 30 req / 60 s |
Degradation ladder: an optional call fails → section omitted plus a warning; the
profile model returns nothing → PROFILE_NOT_VISIBLE; the cookie is rejected →
one re-authentication, then LINKEDIN_SESSION_EXPIRED; repeated failures → the
breaker opens and /readyz reports 503 so an orchestrator can pull the instance
out of rotation.
Observability: structured logs (JSON in production) with a request id on every
line, X-Request-ID / X-Cache / X-Response-Time-Ms headers, /v1/session
diagnostics that never include cookie values, and data.meta.completeness as a
data-quality signal - a drop across many profiles means LinkedIn changed
something, not that members did.
Scaling path (designed, not built): Redis-backed cache and limiter so instances
share both; a session pool of several identities, each with its own cookie
jar, throttle, breaker and egress IP, selected least-recently-used; and a queue
in front of lookups for batch work. VoyagerClient already owns one session,
one throttle and one breaker as a unit, so a pool is a collection of clients
rather than a rewrite. Details in
docs/04-reliability-and-scale.md.
The page at / exists so the API can be evaluated without a terminal: paste a
URL, watch the request happen, read the parsed result, inspect the raw JSON, and
see which upstream calls answered.
- Diagnostics as a first-class tab - every upstream call with status, attempts and latency, plus warnings, so a private contact card is visibly a privacy outcome rather than a bug.
- Coverage meter - 13 segments, one per supported section.
- JSON tab - syntax-highlighted body, copy, download, and the equivalent
curlcommand. - Cache transparency - a hit is labelled with its age;
refreshbypasses it. - Errors that tell you what to do - each code maps to a specific next step.
- Dark and light themes, keyboard shortcut (⌘/Ctrl + K), skeleton loading states, graceful fallbacks when LinkedIn's signed image URLs expire.
No UI framework and no component library: React plus ~23 kB of hand-written CSS,
one accent colour, hairline rules instead of shadows, editorial serif for names,
mono for data. See docs/05-frontend.md.
Every setting is an environment variable, typed in app/config.py and
documented in .env.example. The ones worth reviewing before a deployment:
| Variable | Default | Purpose |
|---|---|---|
LINKEDIN_COOKIE_HEADER |
- | Whole browser Cookie header (recommended) |
LINKEDIN_LI_AT, LINKEDIN_JSESSIONID |
- | The cookie pair, if not using the header |
USER_AGENT |
Chrome 126 on Windows | Should match the browser the cookies came from |
ALLOW_PASSWORD_LOGIN |
false |
Credential login fallback; leave off |
LINKEDIN_PROFILE_QUERY_ID |
- | Optional GraphQL queryId, if you track one |
API_KEYS |
- | Comma separated; unset means the API is open |
CORS_ORIGINS |
* |
Comma separated |
RATE_LIMIT_REQUESTS / _WINDOW_SECONDS |
30 / 60 |
This API's own limit |
CACHE_TTL_SECONDS / CACHE_MAX_ENTRIES |
900 / 512 |
Response cache |
UPSTREAM_MIN_INTERVAL_SECONDS / _JITTER_SECONDS |
1.2 / 0.6 |
Upstream pacing |
UPSTREAM_MAX_CONCURRENCY |
4 |
In-flight upstream calls |
CIRCUIT_BREAKER_THRESHOLD / _COOLDOWN_SECONDS |
4 / 90 |
Breaker |
PROXY_URL |
- | Outbound proxy for egress control |
LOG_FORMAT |
console |
Use json in production |
EXPOSE_RAW_ENDPOINT |
false |
Enables /v1/profile/raw |
uv run pytest -q # 84 tests, no network access
uv run ruff check . # lint
cd web && npm run lint # tsc --noEmit, strict| Suite | Covers |
|---|---|
test_urls.py |
26 URL shapes and rejections |
test_parsers.py |
Payload → schema mapping, field by field, including sparse profiles |
test_client.py |
Headers, retries, 404/410/429/999 mapping, re-auth, breaker, no HTML requests during auth, no duplicate cookies, server-side sign-out detection, Dash path, legacy fallback |
test_service.py |
Cache TTL, LRU, persistence, refresh, coalescing 12 concurrent lookups into one fetch |
test_api.py |
Envelopes, headers, error taxonomy, API key, rate limit, probes, OpenAPI |
respx intercepts LinkedIn at the transport layer, so retries,
re-authentication and the breaker run as real code paths. Fixtures are
synthetic: their shapes come from live captures, their content does not,
so no real person's profile data is committed here.
docs/07-testing.md covers what is deliberately untested.
One container serves both the API and the console. The Dockerfile is two stages
(Node builds the console, Python runs the app), installs with uv sync --frozen,
drops to a non-root user and exposes a /healthz health check.
Blueprints are committed for the two shortest paths to public HTTPS:
# fly.io
fly launch --no-deploy
fly volumes create api_data --size 1
fly secrets set LINKEDIN_COOKIE_HEADER="…" API_KEYS="…"
fly deploy
# Render: render.yaml, Docker runtime, persistent disk, secrets marked sync: falseBoth terminate TLS at the edge; the app runs with --proxy-headers. Keep the
session and cache on a small persistent volume so a redeploy does not
re-authenticate. Operational runbook:
docs/06-deployment.md.
- Visibility is the account's visibility. Contact details usually need a
first-degree connection; out-of-network profiles can return
PROFILE_NOT_VISIBLE. The response says which calls answered. - Session durability is the hard operational problem. A cookie lifted from a browser and replayed from a server is, to LinkedIn, a new device. Mitigations are implemented (full cookie set, rotated cookies retained, CSRF re-seeding, sign-out detection, password login off), but for sustained use an egress IP matching where the cookie was issued matters more than anything in this code.
- One identity, one request budget. The session pool is designed, not built.
- Follower/connection counts and endorsement counts are usually
null- they lived on endpoints LinkedIn retired. - Contact details are usually absent; a birthday often survives, because the decorated profile includes it.
- Sections LinkedIn serves only over GraphQL are not read (featured items,
recommendations, per-position skill tags). The call builder and the normalised
parser exist; no rotating
queryIdis hardcoded. - Grouped positions are flattened - one entry per role, inheriting the group's company details.
- These are internal, unversioned endpoints. They can change without notice;
the blast radius is confined to
endpoints.pyand the parsers, anddata.metamakes a break visible immediately.
The full list, with the reasoning and a prioritised roadmap, is in
docs/08-limitations.md.
Written as the build progressed:
| # | Document |
|---|---|
| 01 | Reverse engineering LinkedIn's profile API |
| 02 | Architecture |
| 03 | Response schema |
| 04 | Reliability and scale |
| 05 | The console |
| 06 | Deployment and operations |
| 07 | Testing |
| 08 | Limitations and roadmap |
Repository conventions and invariants are in AGENTS.md.
Automated collection of LinkedIn profile data is contrary to LinkedIn's User Agreement, regardless of technique. This project was built for a hiring exercise that specified a reverse-engineered, browserless solution. It is rate-limited by default, reads only what the configured account can already see, stores nothing beyond a short-lived cache, and commits no credentials and no real person's profile data. The data belongs to LinkedIn and to the members it describes.
Licensed under the MIT Licence - see LICENSE.
{ "success": true, "meta": { "request_id": "8f2c1d4b9a70", "cached": false, "cache_age_seconds": null, "elapsed_ms": 1180, "upstream_calls": 5 }, "data": { "profile_url": "https://www.linkedin.com/in/adalovelace/", "public_identifier": "adalovelace", "member_urn": "urn:li:member:987654321", "profile_id": "ACoAAB1cEXAMPLE", "first_name": "Ada", "last_name": "Lovelace", "full_name": "Ada Lovelace", "headline": "Principal Engineer at Analytical Engines", "about": "Mathematician turned systems engineer…", "industry": "Computer Software", "pronouns": "She her", "location": { "label": "London Area, United Kingdom", "city": "London", "country_code": "GB", "geo_urn": "urn:li:geo:102257491", "postal_code": "EC1A" }, "profile_picture": { "url": "https://media.licdn.com/dms/image/…/800_800/…", "renditions": [{ "url": "…", "width": 800, "height": 800 }] }, "background_picture": { "url": "…", "renditions": [] }, "network": { "followers": 18422, "connections": null, "distance": null, "following": false }, "contact": { "emails": [], "phone_numbers": [], "websites": [], "birthday": { "month": 12, "day": 10, "label": "10 Dec" } }, "experience": [ { "title": "Principal Engineer", "employment_type": "Full time", "company": { "name": "Analytical Engines", "urn": "urn:li:organization:1441", "universal_name": "analytical-engines", "linkedin_url": "https://www.linkedin.com/company/analytical-engines/", "logo": { "url": "…", "renditions": [] } }, "location": "London, United Kingdom", "description": "Owning the compute platform end to end.", "dates": { "start": { "year": 2021, "month": 3, "iso": "2021-03-01", "label": "Mar 2021" }, "end": null, "is_current": true, "duration_months": 66, "label": "Mar 2021 - Present (5 yrs 6 mos)" } } ], "education": [ { "school": { "name": "University of London" }, "degree": "MSc", "field_of_study": "Mathematics", "grade": "Distinction", "dates": { "label": "2014 - 2016" } } ], "skills": [ { "name": "Distributed Systems", "endorsement_count": 42 } ], "certifications": [ { "name": "Certified Kubernetes Administrator", "authority": "The Linux Foundation", "license_number": "CKA-2023-0917", "dates": { "label": "Sep 2023" } } ], "languages": [ { "name": "English", "proficiency": "Native or bilingual" } ], "projects": [], "publications": [], "honors": [], "volunteering": [], "courses": [], "patents": [], "test_scores": [], "organizations": [], "meta": { "fetched_at": "2026-08-27T09:14:22.481Z", "sources": [ { "endpoint": "dashProfile", "status_code": 200, "ok": true, "elapsed_ms": 412, "attempts": 1 }, { "endpoint": "contactInfo", "status_code": 403, "ok": false, "elapsed_ms": 0, "attempts": 1 } ], "warnings": [ "Call 'contactInfo' unavailable (HTTP 403): contact details are not shared with the querying account" ], "sections_populated": ["experience", "education", "skills"], "completeness": 0.231 } } }