Rish-it/linkedin-profile-api

Reverse-engineered LinkedIn profile API — hits Voyager endpoints directly over HTTP, no browser. FastAPI + curl_cffi Chrome TLS impersonation.

★ 1Forks 0PythonGitHub ↗Compare

README

LinkedIn Profile API

An HTTPS API that takes a LinkedIn profile URL and returns the profile as structured JSON.

It is a reverse-engineered HTTP client. There is no browser, no headless Chrome, no Playwright, no HTML scraping — the service talks directly to LinkedIn's private Voyager JSON API using an authenticated session cookie, exactly the way linkedin.com's own frontend does.

POST /profile  {"url": "https://www.linkedin.com/in/ada-lovelace/"}
   -> { name, headline, location, about, experience[], education[],
        skills[], certifications[], languages[], images }

Contents


Quick start

git clone <this-repo> && cd tross-assignment

python3 -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt

cp .env.example .env      # then paste your own cookies into .env
uvicorn app:app --reload

Then:

curl "http://127.0.0.1:8000/profile?url=https://www.linkedin.com/in/williamhgates/"

.env is gitignored. No credential is ever committed; in production the same two values are set as environment variables on the host.


Getting the session cookies

Voyager only answers an authenticated member session, so the service authenticates as a real LinkedIn account. Two cookies are needed:

Variable What it is
LI_AT The session token. HttpOnly, so document.cookie cannot read it.
JSESSIONID Doubles as the CSRF token. Looks like "ajax:1234567890".

To collect them:

  1. Log in to linkedin.com in a browser.
  2. Open DevTools → Network, reload the page.
  3. Click any linkedin.com request → Request Headers → cookie:.
  4. Copy the li_at= and JSESSIONID= values out of that header.

Keep the surrounding quotes on JSESSIONID. The client strips them for the CSRF header and re-adds them for the cookie jar, because LinkedIn requires the value quoted in one place and bare in the other — an easy thing to get wrong, and the usual cause of a silent 403.


API reference

GET /

One URL, two audiences — the route branches on the Accept header.

A browser (Accept: text/html) gets a demo page: paste a profile URL, see the parsed result rendered, with the raw JSON one click away. It is a single dependency-free static file, so the deployed service is testable without a terminal.

Any API client gets the usage document, unchanged:

{
  "service": "LinkedIn Profile API",
  "usage": "GET /profile?url=https://www.linkedin.com/in/<slug>/"
}

HEAD / returns 200 for platform health probes.

GET /profile?url=<linkedin profile url>

POST /profile — body {"url": "<linkedin profile url>"}

Both accept any form of profile URL and return the same payload:

curl "https://<your-host>/profile?url=https://www.linkedin.com/in/ada-lovelace/"

curl -X POST https://<your-host>/profile \
     -H 'content-type: application/json' \
     -d '{"url":"https://www.linkedin.com/in/ada-lovelace/"}'

Accepted URL shapes — the slug is extracted, everything else ignored:

https://www.linkedin.com/in/ada-lovelace
https://www.linkedin.com/in/ada-lovelace/
https://www.linkedin.com/in/ada-lovelace/?trk=nav&foo=1
https://www.linkedin.com/in/ada-lovelace/details/experience/
https://uk.linkedin.com/in/ada-lovelace
www.linkedin.com/in/ada-lovelace
https://www.linkedin.com/in/jos%C3%A9-garc%C3%ADa      (percent-decoded)

Response 200 OK (real output, produced by the parser from the test fixture):

{
  "public_id": "ada-lovelace",
  "name": "Ada Lovelace",
  "first_name": "Ada",
  "last_name": "Lovelace",
  "headline": "Mathematician | Wrote the first algorithm",
  "location": "London, England, United Kingdom",
  "about": "I write algorithms for machines that do not exist yet.",
  "profile_picture": "https://media.licdn.com/dms/image/pp/800_800/pp.jpg",
  "background_image": "https://media.licdn.com/dms/image/bg/1400_425/bg.jpg",
  "experience": [
    {
      "title": "Analytical Engine Collaborator",
      "company": "Babbage Engineering",
      "location": "London, United Kingdom",
      "description": "Translated Menabrea's memoir; appended Note G.",
      "start": "1842-10",
      "end": "1843-08"
    },
    {
      "title": "Independent Mathematician",
      "company": "Self-employed",
      "location": null,
      "description": null,
      "start": "1833",
      "end": null
    }
  ],
  "education": [
    {
      "school": "University of London",
      "degree": "Private Tuition",
      "field": "Mathematics",
      "start": "1829",
      "end": "1835"
    }
  ],
  "skills": ["Algorithms", "Symbolic Logic"],
  "certifications": [
    {
      "name": "Note G Certification",
      "authority": "Royal Society",
      "url": "https://example.org/note-g",
      "start": "1843-09"
    }
  ],
  "languages": [
    { "name": "French", "proficiency": "PROFESSIONAL_WORKING" }
  ]
}

Every field is nullable and every list may be empty — a profile with no certifications returns "certifications": [], never an error.

Dates are "YYYY-MM" when LinkedIn gives a month and "YYYY" when it only gives a year. An ongoing role has "end": null.

Error responses

Status Meaning
422 The URL is not a LinkedIn /in/ profile URL (company pages, junk, missing param).
500 The server has no LinkedIn credentials configured.
502 LinkedIn rejected the session cookie, or returned an unusable response.
503 LinkedIn rate-limited or challenged the request; retry later.

The distinction between 502 and 503 matters operationally: 502 means rotate the cookie, 503 means slow down.


How it was reverse-engineered

1. Finding the API

Loading a LinkedIn profile with DevTools → Network → Fetch/XHR shows the page is a shell; the content arrives as JSON from linkedin.com/voyager/api/.... Voyager is LinkedIn's internal Rest.li API. Replaying one of those requests in isolation, then stripping it header by header, shows what it actually needs:

  • cookie: li_at=… — the member session.
  • csrf-token: — must equal the JSESSIONID value without quotes.
  • accept: application/vnd.linkedin.normalized+json+2.1 — the flattened response format (see below). Without it you get deeply nested Rest.li.
  • x-restli-protocol-version: 2.0.0.
  • x-li-track — a client fingerprint blob; the server rejects a request claiming to be the web app without it.

Everything else in the browser's request is decoration.

2. The old endpoint is dead

Nearly every LinkedIn scraping tutorial and the popular linkedin-api PyPI package use:

GET /voyager/api/identity/profiles/{public_id}/profileView

That endpoint now returns 410 Gone — LinkedIn retired it. This was the first real finding of the exercise, and the reason a copy-pasted solution cannot work today.

The live replacement is the "dash" (Rest.li 2.0) endpoint:

GET /voyager/api/identity/dash/profiles
      ?q=memberIdentity
      &memberIdentity=<public id>
      &decorationId=com.linkedin.voyager.dash.deco.identity.profile.FullProfileWithEntities-101

The decorationId is the useful part: it is a server-side projection that tells Voyager how much of the object graph to inline. FullProfileWithEntities returns the profile plus positions, educations, skills, certifications and languages in a single response, which is what makes a one-request-per-profile API possible.

3. Reading the response

With the normalized+json accept header the response is:

{ "data": { ... }, "included": [ { "$type": "...Position", ... }, ... ] }

included is a flat pool of every entity referenced anywhere in the graph, each tagged with a $type. So the parser ignores nesting entirely and buckets included by the last segment of $type (Profile, Position, Education, Skill, Certification, Language).

This is deliberate. LinkedIn reshuffles how sections are nested far more often than it renames the entity types, so a type-bucketing parser survives layout changes that a path-walking parser would not.

4. Getting past the anti-abuse layer

A correct, fully authenticated request still gets blocked if it does not look like a browser. Two mechanisms matter:

TLS fingerprinting. LinkedIn fingerprints the TLS ClientHello (JA3). Python's requests/urllib3 produce an OpenSSL fingerprint that no real browser emits, which is detectable before a single HTTP byte is parsed. This project uses curl_cffi with impersonate="chrome", which replays Chrome's exact TLS and HTTP/2 fingerprint. It is still a plain HTTP client — no browser is launched — but at the transport layer it is indistinguishable from Chrome.

Routing cookies. A browser never hits Voyager as its first request. It loads an HTML page, which sets edge-routing cookies (lidc, bcookie). Calling the API cold, without them, produces a self-redirecting 302 loop. The client therefore performs a warmup() — one GET /feed/ — before its first API call, and reuses the session afterwards.

Failure signature. LinkedIn signals a dead session two different ways: a bare 401, or — more confusingly — a 302 carrying Set-Cookie: li_at=delete me; Max-Age=0. The client watches for that across all Set-Cookie headers and raises AuthError (→ 502), distinguishing a dead cookie from ordinary throttling, which gets exponential backoff (2s → 4s → 8s) and then BlockedError (→ 503).


Architecture

Three modules, one job each:

app.py        FastAPI surface: URL -> public id, error mapping to HTTP status
linkedin.py   Voyager HTTP client: auth, warmup, headers, retry, backoff
parse.py      normalized Voyager JSON -> flat profile schema
index.html    demo page served to browsers; no build step, no dependencies

Request flow:

POST /profile
  -> extract_public_id()          regex the /in/<slug>, percent-decode  (422 on miss)
  -> LinkedInClient.warmup()      GET /feed/ for routing cookies, once per client
  -> GET /identity/dash/profiles  with csrf + normalized accept, retry on 302/403/429/999
  -> parse_profile()              bucket `included` by $type, assemble flat JSON
  -> 200

Nothing is cached and no state is stored. Each request builds a client, performs a warmup and one API call, and returns.


Tests

pip install -r requirements-dev.txt
pytest -q
pytest --cov=. --cov-report=term-missing      # coverage

51 tests, 100% line coverage of app.py, linkedin.py and parse.py. The whole suite runs offline in under a second — no network, no LinkedIn account needed, so it is safe to run in CI.

The suite was written test-first, and caught six real defects — four during the build, two more after the first live deployment:

Test Defect found
test_profile_picture_picks_largest_artifact Image picker took the last artifact rather than the largest; LinkedIn does not order them by size.
test_null_date_range_does_not_crash "dateRange": null (present but null, common on ongoing roles) crashed the parser with AttributeError.
test_rejected_cookie_found_among_several_set_cookie_headers Only the first Set-Cookie header was inspected, so a dead cookie was misreported as rate-limiting (503 instead of 502).
test_unexpected_status_raises_linkedin_error An unexpected upstream status leaked a raw transport exception, surfacing as 500 instead of 502.
test_unauthorized_status_raises_auth_error A stale cookie draws a plain 401 rather than the li_at=delete redirect, and was reported as a generic upstream failure instead of "rotate the cookie". Found by deploying with an expired cookie.
test_head_root_is_allowed HEAD / returned 405, so the platform health check logged a failure on every probe.

LinkedIn is stubbed at the HTTP-session boundary, so the client's real logic — header construction, warmup ordering, retry, backoff, error classification — is exercised rather than mocked away.


Deployment

render.yaml deploys to Render's free web tier over HTTPS:

  1. Push this repo to GitHub.
  2. Render → New → Blueprint → select the repo. render.yaml is picked up automatically.
  3. Set LI_AT and JSESSIONID in the Render dashboard (both are declared sync: false, so they are never read from the repo).
  4. Deploy. Render issues an HTTPS subdomain.

To configure a service by hand instead of from the blueprint, the only two settings that matter are the build and start commands:

pip install -r requirements.txt
uvicorn app:app --host 0.0.0.0 --port $PORT

Any host that can run those two lines works — the app has no platform-specific code.


Known limitations

Stated plainly, because they are inherent to the approach rather than oversights:

Session cookies expire. li_at is a browser session token, typically valid for weeks but revoked sooner if LinkedIn sees unusual activity. When it dies the API returns 502 and the cookie must be replaced. There is no way around this: Voyager has no API keys, no OAuth, no service accounts. Any reverse-engineered LinkedIn client has this ceiling.

Rate limits are real and unpublished. LinkedIn throttles aggressively on volume and on request patterns that do not look human. The client backs off and reports 503, but sustained high-volume use will get the account restricted. A production version would need a rotating cookie pool, request pacing, and per-account budgets.

Serving what the logged-in account can see. Results reflect the authenticated member's visibility: connection degree, privacy settings and regional restrictions all change what a profile returns. Two different cookies can legitimately produce different output for the same URL.

location is always null. Confirmed against three live profiles. Every other field returns real data — name, headline, about, both images, experience (including per-role location and dates), education and skills all verified live. The profile's own location is not a plain string on the profile entity the way geoLocationName suggests; it resolves through a URN to a separate entity in included, which the parser does not yet bucket. It is a missing field, not a failing request, and the fix is one more $type bucket.

Certifications and languages are unverified. Both return [] on every profile tested so far, but none of those profiles has either section filled in, so there is no evidence yet distinguishing "correctly empty" from "wrong field name". Skills followed exactly this pattern — empty on two profiles, then 20 entries on the third — so the mapping is probably right, but it is untested rather than proven.

Rare sections are not mapped at all. Volunteering, publications, patents and recommendations are ignored. Each would be one additional $type bucket.

The parser is defensive by construction throughout: an unknown or missing section yields null or [] rather than an error, so a profile shape nobody anticipated degrades instead of failing.

No caching, one profile per request. Each call does a warmup plus one API call. Fine for the intended use, wasteful at volume — the warmup and client should be pooled and reused if throughput ever matters.

Free-tier cold starts. Render's free plan sleeps a service after ~15 minutes of inactivity; the first request after that takes roughly 50 seconds while the instance wakes. Subsequent requests are fast. Any paid tier removes this.

Terms of service. Automated access to LinkedIn using a member session violates LinkedIn's User Agreement, and the account used can be restricted or banned. This project was built as a reverse-engineering exercise. Use a throwaway account, not a primary one.

Contributors

Rish-it

Issues