An HTTPS API that takes a LinkedIn profile URL and returns the profile as structured JSON.
It is a reverse-engineered HTTP client. There is no browser, no headless Chrome, no Playwright, no HTML scraping — the service talks directly to LinkedIn's private Voyager JSON API using an authenticated session cookie, exactly the way linkedin.com's own frontend does.
POST /profile {"url": "https://www.linkedin.com/in/ada-lovelace/"}
-> { name, headline, location, about, experience[], education[],
skills[], certifications[], languages[], images }
- Quick start
- Getting the session cookies
- API reference
- How it was reverse-engineered
- Architecture
- Tests
- Deployment
- Known limitations
git clone <this-repo> && cd tross-assignment
python3 -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
cp .env.example .env # then paste your own cookies into .env
uvicorn app:app --reloadThen:
curl "http://127.0.0.1:8000/profile?url=https://www.linkedin.com/in/williamhgates/".env is gitignored. No credential is ever committed; in production the same
two values are set as environment variables on the host.
Voyager only answers an authenticated member session, so the service authenticates as a real LinkedIn account. Two cookies are needed:
| Variable | What it is |
|---|---|
LI_AT |
The session token. HttpOnly, so document.cookie cannot read it. |
JSESSIONID |
Doubles as the CSRF token. Looks like "ajax:1234567890". |
To collect them:
- Log in to linkedin.com in a browser.
- Open DevTools → Network, reload the page.
- Click any
linkedin.comrequest → Request Headers →cookie:. - Copy the
li_at=andJSESSIONID=values out of that header.
Keep the surrounding quotes on JSESSIONID. The client strips them for the
CSRF header and re-adds them for the cookie jar, because LinkedIn requires the
value quoted in one place and bare in the other — an easy thing to get wrong,
and the usual cause of a silent 403.
One URL, two audiences — the route branches on the Accept header.
A browser (Accept: text/html) gets a demo page: paste a profile URL, see
the parsed result rendered, with the raw JSON one click away. It is a single
dependency-free static file, so the deployed service is testable without a
terminal.
Any API client gets the usage document, unchanged:
{
"service": "LinkedIn Profile API",
"usage": "GET /profile?url=https://www.linkedin.com/in/<slug>/"
}HEAD / returns 200 for platform health probes.
Both accept any form of profile URL and return the same payload:
curl "https://<your-host>/profile?url=https://www.linkedin.com/in/ada-lovelace/"
curl -X POST https://<your-host>/profile \
-H 'content-type: application/json' \
-d '{"url":"https://www.linkedin.com/in/ada-lovelace/"}'Accepted URL shapes — the slug is extracted, everything else ignored:
https://www.linkedin.com/in/ada-lovelace
https://www.linkedin.com/in/ada-lovelace/
https://www.linkedin.com/in/ada-lovelace/?trk=nav&foo=1
https://www.linkedin.com/in/ada-lovelace/details/experience/
https://uk.linkedin.com/in/ada-lovelace
www.linkedin.com/in/ada-lovelace
https://www.linkedin.com/in/jos%C3%A9-garc%C3%ADa (percent-decoded)
Response 200 OK (real output, produced by the parser from the test fixture):
{
"public_id": "ada-lovelace",
"name": "Ada Lovelace",
"first_name": "Ada",
"last_name": "Lovelace",
"headline": "Mathematician | Wrote the first algorithm",
"location": "London, England, United Kingdom",
"about": "I write algorithms for machines that do not exist yet.",
"profile_picture": "https://media.licdn.com/dms/image/pp/800_800/pp.jpg",
"background_image": "https://media.licdn.com/dms/image/bg/1400_425/bg.jpg",
"experience": [
{
"title": "Analytical Engine Collaborator",
"company": "Babbage Engineering",
"location": "London, United Kingdom",
"description": "Translated Menabrea's memoir; appended Note G.",
"start": "1842-10",
"end": "1843-08"
},
{
"title": "Independent Mathematician",
"company": "Self-employed",
"location": null,
"description": null,
"start": "1833",
"end": null
}
],
"education": [
{
"school": "University of London",
"degree": "Private Tuition",
"field": "Mathematics",
"start": "1829",
"end": "1835"
}
],
"skills": ["Algorithms", "Symbolic Logic"],
"certifications": [
{
"name": "Note G Certification",
"authority": "Royal Society",
"url": "https://example.org/note-g",
"start": "1843-09"
}
],
"languages": [
{ "name": "French", "proficiency": "PROFESSIONAL_WORKING" }
]
}Every field is nullable and every list may be empty — a profile with no
certifications returns "certifications": [], never an error.
Dates are "YYYY-MM" when LinkedIn gives a month and "YYYY" when it only
gives a year. An ongoing role has "end": null.
Error responses
| Status | Meaning |
|---|---|
422 |
The URL is not a LinkedIn /in/ profile URL (company pages, junk, missing param). |
500 |
The server has no LinkedIn credentials configured. |
502 |
LinkedIn rejected the session cookie, or returned an unusable response. |
503 |
LinkedIn rate-limited or challenged the request; retry later. |
The distinction between 502 and 503 matters operationally: 502 means
rotate the cookie, 503 means slow down.
Loading a LinkedIn profile with DevTools → Network → Fetch/XHR shows the page
is a shell; the content arrives as JSON from linkedin.com/voyager/api/....
Voyager is LinkedIn's internal Rest.li API. Replaying one of those requests in
isolation, then stripping it header by header, shows what it actually needs:
cookie: li_at=…— the member session.csrf-token:— must equal theJSESSIONIDvalue without quotes.accept: application/vnd.linkedin.normalized+json+2.1— the flattened response format (see below). Without it you get deeply nested Rest.li.x-restli-protocol-version: 2.0.0.x-li-track— a client fingerprint blob; the server rejects a request claiming to be the web app without it.
Everything else in the browser's request is decoration.
Nearly every LinkedIn scraping tutorial and the popular linkedin-api PyPI
package use:
GET /voyager/api/identity/profiles/{public_id}/profileView
That endpoint now returns 410 Gone — LinkedIn retired it. This was the
first real finding of the exercise, and the reason a copy-pasted solution
cannot work today.
The live replacement is the "dash" (Rest.li 2.0) endpoint:
GET /voyager/api/identity/dash/profiles
?q=memberIdentity
&memberIdentity=<public id>
&decorationId=com.linkedin.voyager.dash.deco.identity.profile.FullProfileWithEntities-101
The decorationId is the useful part: it is a server-side projection that tells
Voyager how much of the object graph to inline. FullProfileWithEntities returns
the profile plus positions, educations, skills, certifications and languages
in a single response, which is what makes a one-request-per-profile API possible.
With the normalized+json accept header the response is:
{ "data": { ... }, "included": [ { "$type": "...Position", ... }, ... ] }included is a flat pool of every entity referenced anywhere in the graph, each
tagged with a $type. So the parser ignores nesting entirely and buckets
included by the last segment of $type (Profile, Position, Education,
Skill, Certification, Language).
This is deliberate. LinkedIn reshuffles how sections are nested far more often than it renames the entity types, so a type-bucketing parser survives layout changes that a path-walking parser would not.
A correct, fully authenticated request still gets blocked if it does not look like a browser. Two mechanisms matter:
TLS fingerprinting. LinkedIn fingerprints the TLS ClientHello (JA3). Python's
requests/urllib3 produce an OpenSSL fingerprint that no real browser emits,
which is detectable before a single HTTP byte is parsed. This project uses
curl_cffi with impersonate="chrome",
which replays Chrome's exact TLS and HTTP/2 fingerprint. It is still a plain HTTP
client — no browser is launched — but at the transport layer it is
indistinguishable from Chrome.
Routing cookies. A browser never hits Voyager as its first request. It loads
an HTML page, which sets edge-routing cookies (lidc, bcookie). Calling the
API cold, without them, produces a self-redirecting 302 loop. The client
therefore performs a warmup() — one GET /feed/ — before its first API call,
and reuses the session afterwards.
Failure signature. LinkedIn signals a dead session two different ways: a bare
401, or — more confusingly — a 302 carrying Set-Cookie: li_at=delete me; Max-Age=0. The
client watches for that across all Set-Cookie headers and raises AuthError
(→ 502), distinguishing a dead cookie from ordinary throttling, which gets
exponential backoff (2s → 4s → 8s) and then BlockedError (→ 503).
Three modules, one job each:
app.py FastAPI surface: URL -> public id, error mapping to HTTP status
linkedin.py Voyager HTTP client: auth, warmup, headers, retry, backoff
parse.py normalized Voyager JSON -> flat profile schema
index.html demo page served to browsers; no build step, no dependencies
Request flow:
POST /profile
-> extract_public_id() regex the /in/<slug>, percent-decode (422 on miss)
-> LinkedInClient.warmup() GET /feed/ for routing cookies, once per client
-> GET /identity/dash/profiles with csrf + normalized accept, retry on 302/403/429/999
-> parse_profile() bucket `included` by $type, assemble flat JSON
-> 200
Nothing is cached and no state is stored. Each request builds a client, performs a warmup and one API call, and returns.
pip install -r requirements-dev.txt
pytest -q
pytest --cov=. --cov-report=term-missing # coverage51 tests, 100% line coverage of app.py, linkedin.py and parse.py. The
whole suite runs offline in under a second — no network, no LinkedIn account
needed, so it is safe to run in CI.
The suite was written test-first, and caught six real defects — four during the build, two more after the first live deployment:
| Test | Defect found |
|---|---|
test_profile_picture_picks_largest_artifact |
Image picker took the last artifact rather than the largest; LinkedIn does not order them by size. |
test_null_date_range_does_not_crash |
"dateRange": null (present but null, common on ongoing roles) crashed the parser with AttributeError. |
test_rejected_cookie_found_among_several_set_cookie_headers |
Only the first Set-Cookie header was inspected, so a dead cookie was misreported as rate-limiting (503 instead of 502). |
test_unexpected_status_raises_linkedin_error |
An unexpected upstream status leaked a raw transport exception, surfacing as 500 instead of 502. |
test_unauthorized_status_raises_auth_error |
A stale cookie draws a plain 401 rather than the li_at=delete redirect, and was reported as a generic upstream failure instead of "rotate the cookie". Found by deploying with an expired cookie. |
test_head_root_is_allowed |
HEAD / returned 405, so the platform health check logged a failure on every probe. |
LinkedIn is stubbed at the HTTP-session boundary, so the client's real logic — header construction, warmup ordering, retry, backoff, error classification — is exercised rather than mocked away.
render.yaml deploys to Render's free web tier over HTTPS:
- Push this repo to GitHub.
- Render → New → Blueprint → select the repo.
render.yamlis picked up automatically. - Set
LI_ATandJSESSIONIDin the Render dashboard (both are declaredsync: false, so they are never read from the repo). - Deploy. Render issues an HTTPS subdomain.
To configure a service by hand instead of from the blueprint, the only two settings that matter are the build and start commands:
pip install -r requirements.txt
uvicorn app:app --host 0.0.0.0 --port $PORT
Any host that can run those two lines works — the app has no platform-specific code.
Stated plainly, because they are inherent to the approach rather than oversights:
Session cookies expire. li_at is a browser session token, typically valid
for weeks but revoked sooner if LinkedIn sees unusual activity. When it dies the
API returns 502 and the cookie must be replaced. There is no way around this:
Voyager has no API keys, no OAuth, no service accounts. Any reverse-engineered
LinkedIn client has this ceiling.
Rate limits are real and unpublished. LinkedIn throttles aggressively on
volume and on request patterns that do not look human. The client backs off and
reports 503, but sustained high-volume use will get the account restricted. A
production version would need a rotating cookie pool, request pacing, and
per-account budgets.
Serving what the logged-in account can see. Results reflect the authenticated member's visibility: connection degree, privacy settings and regional restrictions all change what a profile returns. Two different cookies can legitimately produce different output for the same URL.
location is always null. Confirmed against three live profiles. Every
other field returns real data — name, headline, about, both images, experience
(including per-role location and dates), education and skills all verified
live. The profile's own location is not a plain string on the profile entity
the way geoLocationName suggests; it resolves through a URN to a separate
entity in included, which the parser does not yet bucket. It is a missing
field, not a failing request, and the fix is one more $type bucket.
Certifications and languages are unverified. Both return [] on every
profile tested so far, but none of those profiles has either section filled in,
so there is no evidence yet distinguishing "correctly empty" from "wrong field
name". Skills followed exactly this pattern — empty on two profiles, then 20
entries on the third — so the mapping is probably right, but it is untested
rather than proven.
Rare sections are not mapped at all. Volunteering, publications, patents and
recommendations are ignored. Each would be one additional $type bucket.
The parser is defensive by construction throughout: an unknown or missing
section yields null or [] rather than an error, so a profile shape nobody
anticipated degrades instead of failing.
No caching, one profile per request. Each call does a warmup plus one API call. Fine for the intended use, wasteful at volume — the warmup and client should be pooled and reused if throughput ever matters.
Free-tier cold starts. Render's free plan sleeps a service after ~15 minutes of inactivity; the first request after that takes roughly 50 seconds while the instance wakes. Subsequent requests are fast. Any paid tier removes this.
Terms of service. Automated access to LinkedIn using a member session violates LinkedIn's User Agreement, and the account used can be restricted or banned. This project was built as a reverse-engineering exercise. Use a throwaway account, not a primary one.