priyank766/Tross

★ 0Forks 0PythonGitHub ↗Compare

README

LinkedIn Profile API

A hosted read API that accepts a LinkedIn profile URL and returns the profile page as structured JSON: identity, about, experience, education, skills, certifications, languages, projects, publications, honours, volunteering, courses, patents, test scores, images and provenance.

It is purely reverse-engineered and browserless. The service speaks LinkedIn's own internal web API (Voyager) over HTTPS with a member session cookie held server-side. There is no headless browser, no automation driver and no HTML parsing anywhere in the request path.


Contents


Quick start

Requires uv and, for the console, Node 20+.

git clone <this-repo> && cd linkedin-profile-api
cp .env.example .env          # then fill in the cookie header, see Credentials
uv sync --extra dev           # creates .venv and installs dependencies
cd web && npm install && npm run build && cd ..   # builds the console into app/static
uv run uvicorn app.main:app --port 8000

Open http://127.0.0.1:8000 for the console, or call the API directly:

curl -s "http://127.0.0.1:8000/v1/profile?url=https://www.linkedin.com/in/williamhgates/" | jq

One profile, straight from the terminal, without starting the server:

uv run linkedin-api "https://www.linkedin.com/in/williamhgates/"
uv run linkedin-api "williamhgates" --meta-only     # provenance and coverage only

With Docker (builds the console and the API into one image):

docker build -t linkedin-profile-api .
docker run -p 8000:8000 --env-file .env linkedin-profile-api

Credentials

The service authenticates as a real LinkedIn member. Nothing is committed: .env is git-ignored, .env.example documents every variable, and platform deployments take secrets out of band.

Recommended - the whole cookie header. In a signed-in browser: DevTools (F12) → Network → reload → click the first document request → Request Headers → copy the value of cookie. Then, as a single line:

LINKEDIN_COOKIE_HEADER=li_at=AQEDA…; JSESSIONID="ajax:12345…"; bcookie="v=2&…"; bscookie="v=1&…"; lidc="b=…"
USER_AGENT=<paste navigator.userAgent from the same browser>

Minimum - the cookie pair. li_at and JSESSIONID, copied at the same time from DevTools → Application → Cookies → https://www.linkedin.com:

LINKEDIN_LI_AT=AQEDA…
LINKEDIN_JSESSIONID=ajax:12345…

Why the whole header is worth the extra step: LinkedIn reads device identity from bcookie/bscookie and datacenter routing from lidc. An auth token arriving without them looks like a different device using a copied token, and LinkedIn responds by expiring the session server-side (observed live: a 302 to the same URL plus Set-Cookie: li_at=delete me). This service detects that signal and reports LINKEDIN_SESSION_EXPIRED with instructions rather than looping.

Credential login (/uas/authenticate) is implemented but off by default (ALLOW_PASSWORD_LOGIN=false): a failed attempt can invalidate a working cookie session and can put the account into a verification challenge.

Check the session at any time:

curl -s localhost:8000/readyz | jq          # liveness of the upstream session
curl -s localhost:8000/v1/session | jq      # source, age, breaker and cache state

API reference

Base path /v1. Every response is JSON and carries X-Request-ID, X-Cache (HIT/MISS) and X-Response-Time-Ms.

GET /v1/profile

Parameter Type Default Notes
url string, required A LinkedIn profile URL, or the bare vanity name
refresh boolean false Bypass the cache and refetch upstream

Accepted inputs - all of these resolve to the same profile:

https://www.linkedin.com/in/williamhgates/
https://www.linkedin.com/in/williamhgates?trk=feed&originalSubdomain=us
https://in.linkedin.com/in/williamhgates
https://m.linkedin.com/in/williamhgates/
linkedin.com/in/williamhgates
williamhgates
https://www.linkedin.com/in/ACoAAA8BYqEB…      (obfuscated member id)
curl -s "$HOST/v1/profile?url=williamhgates" -H "X-API-Key: $KEY"

POST /v1/profile

curl -s "$HOST/v1/profile" -H 'content-type: application/json' \
  -d '{"url": "https://www.linkedin.com/in/williamhgates/", "refresh": false}'

Other endpoints

Endpoint Purpose
GET /healthz Liveness: version, environment, uptime
GET /readyz Readiness: verifies the LinkedIn session, 503 when it is unusable
GET /v1/session Session source and age, circuit breaker state, cache hit rate. Never returns cookie values
GET /v1/profile/raw?url= Unparsed upstream payloads, for debugging. Disabled unless EXPOSE_RAW_ENDPOINT=true
GET /docs, /redoc, /openapi.json Generated API documentation

Errors

Every failure uses one shape, with a stable code to branch on:

{
  "success": false,
  "request_id": "8f2c1d4b9a70",
  "error": {
    "code": "PROFILE_NOT_VISIBLE",
    "message": "LinkedIn returned an empty profile projection for this member.",
    "details": { "identifier": "someone" }
  }
}
Code HTTP Meaning
INVALID_PROFILE_URL 422 Not a /in/ member URL (company, school, or malformed)
VALIDATION_ERROR 422 Missing or unknown request field
PROFILE_NOT_FOUND 404 No such member
PROFILE_NOT_VISIBLE 403 Member exists, details withheld from this account
API_KEY_MISSING / API_KEY_INVALID 401 / 403 This API's own authentication
RATE_LIMITED 429 This API's rate limit; Retry-After set
LINKEDIN_RATE_LIMITED 429 LinkedIn throttled us (429 or its non-standard 999)
LINKEDIN_SESSION_EXPIRED 503 Cookie no longer valid, or signed out upstream
LINKEDIN_CHALLENGE_REQUIRED 503 Account needs human verification
LINKEDIN_AUTH_FAILED 503 No usable session configured
UPSTREAM_CIRCUIT_OPEN 503 Breaker paused upstream calls; retry_after_seconds included
LINKEDIN_UNAVAILABLE 502 Unexpected upstream response

Authentication and limits

Set API_KEYS=key1,key2 to require X-API-Key: <key> (or Authorization: Bearer <key>). Left unset, the API is open, which is convenient locally. The default rate limit is 30 requests per 60 seconds per key or client address.


Response schema

Designed for consumers rather than mirroring LinkedIn's internals:

  • snake_case keys, explicit nulls instead of missing keys, and lists that are always lists.
  • Dates split into parts plus a padded iso string and a human label; duration_months is null when the month is unknown rather than inventing precision.
  • Images as rendition sets - url is the largest, renditions[] carries every size with dimensions.
  • Public URNs (urn:li:organization:1441, not urn:li:fs_miniCompany:1441).
  • Humanised enums (Native or bilingual, not NATIVE_OR_BILINGUAL).
  • Self-describing completeness - see data.meta.
{
  "success": true,
  "meta": {
    "request_id": "8f2c1d4b9a70",
    "cached": false,
    "cache_age_seconds": null,
    "elapsed_ms": 1180,
    "upstream_calls": 5
  },
  "data": {
    "profile_url": "https://www.linkedin.com/in/adalovelace/",
    "public_identifier": "adalovelace",
    "member_urn": "urn:li:member:987654321",
    "profile_id": "ACoAAB1cEXAMPLE",
    "first_name": "Ada",
    "last_name": "Lovelace",
    "full_name": "Ada Lovelace",
    "headline": "Principal Engineer at Analytical Engines",
    "about": "Mathematician turned systems engineer…",
    "industry": "Computer Software",
    "pronouns": "She her",
    "location": {
      "label": "London Area, United Kingdom",
      "city": "London",
      "country_code": "GB",
      "geo_urn": "urn:li:geo:102257491",
      "postal_code": "EC1A"
    },
    "profile_picture": {
      "url": "https://media.licdn.com/dms/image/…/800_800/…",
      "renditions": [{ "url": "…", "width": 800, "height": 800 }]
    },
    "background_picture": { "url": "…", "renditions": [] },
    "network": { "followers": 18422, "connections": null, "distance": null, "following": false },
    "contact": { "emails": [], "phone_numbers": [], "websites": [], "birthday": { "month": 12, "day": 10, "label": "10 Dec" } },
    "experience": [
      {
        "title": "Principal Engineer",
        "employment_type": "Full time",
        "company": {
          "name": "Analytical Engines",
          "urn": "urn:li:organization:1441",
          "universal_name": "analytical-engines",
          "linkedin_url": "https://www.linkedin.com/company/analytical-engines/",
          "logo": { "url": "…", "renditions": [] }
        },
        "location": "London, United Kingdom",
        "description": "Owning the compute platform end to end.",
        "dates": {
          "start": { "year": 2021, "month": 3, "iso": "2021-03-01", "label": "Mar 2021" },
          "end": null,
          "is_current": true,
          "duration_months": 66,
          "label": "Mar 2021 - Present (5 yrs 6 mos)"
        }
      }
    ],
    "education": [ { "school": { "name": "University of London" }, "degree": "MSc", "field_of_study": "Mathematics", "grade": "Distinction", "dates": { "label": "2014 - 2016" } } ],
    "skills": [ { "name": "Distributed Systems", "endorsement_count": 42 } ],
    "certifications": [ { "name": "Certified Kubernetes Administrator", "authority": "The Linux Foundation", "license_number": "CKA-2023-0917", "dates": { "label": "Sep 2023" } } ],
    "languages": [ { "name": "English", "proficiency": "Native or bilingual" } ],
    "projects": [], "publications": [], "honors": [], "volunteering": [],
    "courses": [], "patents": [], "test_scores": [], "organizations": [],
    "meta": {
      "fetched_at": "2026-08-27T09:14:22.481Z",
      "sources": [
        { "endpoint": "dashProfile", "status_code": 200, "ok": true, "elapsed_ms": 412, "attempts": 1 },
        { "endpoint": "contactInfo", "status_code": 403, "ok": false, "elapsed_ms": 0, "attempts": 1 }
      ],
      "warnings": [
        "Call 'contactInfo' unavailable (HTTP 403): contact details are not shared with the querying account"
      ],
      "sections_populated": ["experience", "education", "skills"],
      "completeness": 0.231
    }
  }
}

data.meta is the difference between "the API returned nothing for skills" and "this member lists no skills". A full example lives in docs/example-response.json; the authoritative definition is /openapi.json.


Approach

1. Which surface to call. The official APIs are partner-gated and cannot read arbitrary profiles. A browser was ruled out by the brief and is structurally worse (seconds and hundreds of megabytes per profile, markup that changes weekly). Server-rendered HTML is a truncated preview when logged out and an Ember shell when logged in. What remains is the API the LinkedIn web app itself calls: Voyager, at https://www.linkedin.com/voyager/api/….

2. Authentication is two cookies. li_at is the session; JSESSIONID ("ajax:…") must be echoed in a csrf-token header on every call. The service takes them from configuration, validates them with one call to /voyager/api/me, and persists the jar so a restart does not re-authenticate. Repeated logins are the fastest way to get an account flagged.

3. The documented endpoint is gone. identity/profiles/{id}/profileView - the call every open-source LinkedIn client is built on - now answers HTTP 410 Gone, and /profileContactInfo, /skills and /networkinfo answer 403. This was established by probing with a live session, not assumed.

4. The current model is the Dash collection.

GET /voyager/api/identity/dash/profiles
      ?q=memberIdentity
      &memberIdentity=<vanity>
      &decorationId=com.linkedin.voyager.dash.deco.identity.profile.FullProfileWithEntities-101

One call returns the member record plus every section as an embedded tree (profilePositionGroups, profileEducations, profileSkills, …), with companies and schools inlined and images behind displayImageReference. Unlike LinkedIn's GraphQL endpoints it needs no queryId hash, so it does not break on their release schedule.

5. Strategy chain, not a single call. LinkedIn keeps more than one model live, so the repository tries the decorated Dash collection, then an undecorated variant, then GraphQL if a queryId is configured, then the retired legacy projection - stopping at the first that yields a populated profile. Both envelopes (embedded tree, and GraphQL's normalised data + included entity graph) are reduced to one intermediate form, so a single set of pure mappers produces the public schema either way.

6. Enrichment degrades. Contact card, paginated skills and network counts run concurrently and are allowed to fail: a private contact card produces a warning in data.meta.warnings, not a 5xx.

7. Behave like a browser, or lose the session. Requests carry the header set the web client sends (CSRF token, x-restli-protocol-version, x-li-track, a fresh x-li-page-instance UUID, the profile's own referer). Cookies live in a single jar - passing them per request as well made httpx emit two JSESSIONID values, which LinkedIn answers with CSRF check failed. Rotated cookies are absorbed back. No HTML page is ever fetched when a cookie pair is supplied.

Full reasoning, including what was tried and rejected, is in docs/01-reverse-engineering.md.


Architecture

Dependencies point one way: HTTP → services → LinkedIn integration. Nothing below the API layer knows FastAPI exists, and the parsers are pure functions.

app/
├─ api/                 HTTP surface
│  ├─ routes/           profile lookups, health, readiness, session diagnostics
│  ├─ errors.py         error taxonomy and handlers (one envelope for every failure)
│  ├─ security.py       API key authentication (constant-time)
│  └─ ratelimit.py      fixed-window limiter
├─ services/
│  ├─ profile_service.py cache lookup, request coalescing, response metadata
│  └─ cache.py          TTL + LRU cache with optional disk persistence
├─ linkedin/            the reverse-engineered integration
│  ├─ endpoints.py      every upstream URL, in one auditable file
│  ├─ auth.py           cookie and credential flows, header set, sign-out detection
│  ├─ session.py        session value object, atomic persistence
│  ├─ client.py         HTTP/2 pool, retries, re-auth, status → exception mapping
│  ├─ throttle.py       pacing with jitter, concurrency ceiling, circuit breaker
│  ├─ repository.py     strategy chain + concurrent enrichment
│  ├─ normalized.py     resolver for the GraphQL entity-graph envelope
│  └─ parsers/          pure mappers → the published schema
│     ├─ dash.py        current model (embedded and normalised envelopes)
│     ├─ profile.py     retired legacy projection
│     ├─ contact.py     contact card
│     ├─ draft.py       shared intermediate representation
│     └─ assembler.py   draft → public document, with provenance
├─ models/              response schema (Pydantic v2)
├─ config.py            typed settings, the only place env vars are read
└─ static/              built console (generated)

Request lifecycle:

GET /v1/profile?url=…
  ├─ middleware: request id, timing, structured log
  ├─ API key + rate limit
  ├─ parse and canonicalise the URL           → 422 before any network call
  ├─ cache lookup (canonical identifier)      → HIT: return, X-Cache: HIT
  ├─ single-flight guard                      → duplicate concurrent lookups share one fetch
  ├─ strategy chain                           → dashProfile → dashProfileMinimal → graphql → profileView
  ├─ enrichment (concurrent, failures degrade to warnings)
  ├─ pure parsers → draft → public document
  └─ cache store → envelope

Reliability and scale

The scarce resource is not CPU - it is the LinkedIn account's request budget.

Mechanism Default
Response cache (TTL + LRU, optional disk persistence) 900 s, 512 entries
Single-flight coalescing (N concurrent lookups → 1 upstream fetch) always on
Minimum interval + jitter between upstream calls 1.2 s + up to 0.6 s
Concurrency ceiling 4 in flight
Retries with exponential backoff (429/5xx only; 410 never retried) 3 attempts
Single re-authentication on a rejected session per call
Circuit breaker (closed → open → half-open trial) 4 failures, 90 s cooldown
Session persistence across restarts on
This API's own rate limit 30 req / 60 s

Degradation ladder: an optional call fails → section omitted plus a warning; the profile model returns nothing → PROFILE_NOT_VISIBLE; the cookie is rejected → one re-authentication, then LINKEDIN_SESSION_EXPIRED; repeated failures → the breaker opens and /readyz reports 503 so an orchestrator can pull the instance out of rotation.

Observability: structured logs (JSON in production) with a request id on every line, X-Request-ID / X-Cache / X-Response-Time-Ms headers, /v1/session diagnostics that never include cookie values, and data.meta.completeness as a data-quality signal - a drop across many profiles means LinkedIn changed something, not that members did.

Scaling path (designed, not built): Redis-backed cache and limiter so instances share both; a session pool of several identities, each with its own cookie jar, throttle, breaker and egress IP, selected least-recently-used; and a queue in front of lookups for batch work. VoyagerClient already owns one session, one throttle and one breaker as a unit, so a pool is a collection of clients rather than a rewrite. Details in docs/04-reliability-and-scale.md.


The console

The page at / exists so the API can be evaluated without a terminal: paste a URL, watch the request happen, read the parsed result, inspect the raw JSON, and see which upstream calls answered.

  • Diagnostics as a first-class tab - every upstream call with status, attempts and latency, plus warnings, so a private contact card is visibly a privacy outcome rather than a bug.
  • Coverage meter - 13 segments, one per supported section.
  • JSON tab - syntax-highlighted body, copy, download, and the equivalent curl command.
  • Cache transparency - a hit is labelled with its age; refresh bypasses it.
  • Errors that tell you what to do - each code maps to a specific next step.
  • Dark and light themes, keyboard shortcut (⌘/Ctrl + K), skeleton loading states, graceful fallbacks when LinkedIn's signed image URLs expire.

No UI framework and no component library: React plus ~23 kB of hand-written CSS, one accent colour, hairline rules instead of shadows, editorial serif for names, mono for data. See docs/05-frontend.md.


Configuration

Every setting is an environment variable, typed in app/config.py and documented in .env.example. The ones worth reviewing before a deployment:

Variable Default Purpose
LINKEDIN_COOKIE_HEADER - Whole browser Cookie header (recommended)
LINKEDIN_LI_AT, LINKEDIN_JSESSIONID - The cookie pair, if not using the header
USER_AGENT Chrome 126 on Windows Should match the browser the cookies came from
ALLOW_PASSWORD_LOGIN false Credential login fallback; leave off
LINKEDIN_PROFILE_QUERY_ID - Optional GraphQL queryId, if you track one
API_KEYS - Comma separated; unset means the API is open
CORS_ORIGINS * Comma separated
RATE_LIMIT_REQUESTS / _WINDOW_SECONDS 30 / 60 This API's own limit
CACHE_TTL_SECONDS / CACHE_MAX_ENTRIES 900 / 512 Response cache
UPSTREAM_MIN_INTERVAL_SECONDS / _JITTER_SECONDS 1.2 / 0.6 Upstream pacing
UPSTREAM_MAX_CONCURRENCY 4 In-flight upstream calls
CIRCUIT_BREAKER_THRESHOLD / _COOLDOWN_SECONDS 4 / 90 Breaker
PROXY_URL - Outbound proxy for egress control
LOG_FORMAT console Use json in production
EXPOSE_RAW_ENDPOINT false Enables /v1/profile/raw

Testing

uv run pytest -q          # 84 tests, no network access
uv run ruff check .       # lint
cd web && npm run lint    # tsc --noEmit, strict
Suite Covers
test_urls.py 26 URL shapes and rejections
test_parsers.py Payload → schema mapping, field by field, including sparse profiles
test_client.py Headers, retries, 404/410/429/999 mapping, re-auth, breaker, no HTML requests during auth, no duplicate cookies, server-side sign-out detection, Dash path, legacy fallback
test_service.py Cache TTL, LRU, persistence, refresh, coalescing 12 concurrent lookups into one fetch
test_api.py Envelopes, headers, error taxonomy, API key, rate limit, probes, OpenAPI

respx intercepts LinkedIn at the transport layer, so retries, re-authentication and the breaker run as real code paths. Fixtures are synthetic: their shapes come from live captures, their content does not, so no real person's profile data is committed here. docs/07-testing.md covers what is deliberately untested.


Deployment

One container serves both the API and the console. The Dockerfile is two stages (Node builds the console, Python runs the app), installs with uv sync --frozen, drops to a non-root user and exposes a /healthz health check.

Blueprints are committed for the two shortest paths to public HTTPS:

# fly.io
fly launch --no-deploy
fly volumes create api_data --size 1
fly secrets set LINKEDIN_COOKIE_HEADER="…" API_KEYS="…"
fly deploy

# Render: render.yaml, Docker runtime, persistent disk, secrets marked sync: false

Both terminate TLS at the edge; the app runs with --proxy-headers. Keep the session and cache on a small persistent volume so a redeploy does not re-authenticate. Operational runbook: docs/06-deployment.md.


Known limitations

  • Visibility is the account's visibility. Contact details usually need a first-degree connection; out-of-network profiles can return PROFILE_NOT_VISIBLE. The response says which calls answered.
  • Session durability is the hard operational problem. A cookie lifted from a browser and replayed from a server is, to LinkedIn, a new device. Mitigations are implemented (full cookie set, rotated cookies retained, CSRF re-seeding, sign-out detection, password login off), but for sustained use an egress IP matching where the cookie was issued matters more than anything in this code.
  • One identity, one request budget. The session pool is designed, not built.
  • Follower/connection counts and endorsement counts are usually null - they lived on endpoints LinkedIn retired.
  • Contact details are usually absent; a birthday often survives, because the decorated profile includes it.
  • Sections LinkedIn serves only over GraphQL are not read (featured items, recommendations, per-position skill tags). The call builder and the normalised parser exist; no rotating queryId is hardcoded.
  • Grouped positions are flattened - one entry per role, inheriting the group's company details.
  • These are internal, unversioned endpoints. They can change without notice; the blast radius is confined to endpoints.py and the parsers, and data.meta makes a break visible immediately.

The full list, with the reasoning and a prioritised roadmap, is in docs/08-limitations.md.


Engineering notes

Written as the build progressed:

# Document
01 Reverse engineering LinkedIn's profile API
02 Architecture
03 Response schema
04 Reliability and scale
05 The console
06 Deployment and operations
07 Testing
08 Limitations and roadmap

Repository conventions and invariants are in AGENTS.md.


Responsible use

Automated collection of LinkedIn profile data is contrary to LinkedIn's User Agreement, regardless of technique. This project was built for a hiring exercise that specified a reverse-engineered, browserless solution. It is rate-limited by default, reads only what the configured account can already see, stores nothing beyond a short-lived cache, and commits no credentials and no real person's profile data. The data belongs to LinkedIn and to the members it describes.

Licensed under the MIT Licence - see LICENSE.

Contributors

priyank766

Issues