yt-transcripts.tjwpier.uk — YouTube transcript-derived views (guide for LLM agents and humans) ====================================================================================================== WHAT THIS IS One read-only HTTP API. Give it a YouTube video (11-char id or any YouTube URL) and get back one of three text views. No auth, plain GET, responses are text. Errors are JSON: {"detail": "..."}. ENDPOINTS GET //transcript -> text/plain full English caption transcript, joined into one block of text GET //transcript-with-timestamps -> text/plain same transcript, one caption cue per line prefixed [m:ss] / [h:mm:ss] (start time) GET //summary -> text/markdown 4-10 numbered bullets (+ optional Key Takeaways), LLM-generated from the FULL transcript GET //semantic -> text/plain deterministic digest, no LLM: FORM / TOPICS / ENTITIES / NUMBERS / ARC / CENTRAL CLAIMS GET /health -> JSON service status, versions, cache counts GET / (also /help, /llms.txt) this page Legacy aliases, still served (all return the transcript): GET / GET / GET /?url= INPUT = 11-character YouTube video id (e.g. nJMtHxgub2o) or a full URL: youtube.com/watch?v=..., youtu.be/..., /shorts/, /embed/, /live/. Prefer the bare id in the path. A pasted URL works unencoded too, with or without a view suffix: /https://www.youtube.com/watch?v=nJMtHxgub2o/summary /https://youtu.be/nJMtHxgub2o/semantic WHICH VIEW TO USE Ranking / triaging many videos ........ /semantic first (deterministic, ~300 tokens, <1 s), then /summary for the shortlist (~500 tokens, LLM quality). "Is this worth watching / what's in it" /summary. Quotes, details, Q&A about the content /transcript — can be 10k-500k chars; read x-transcript-chars (HEAD or first response) before loading it into context. "Where in the video is X" / chapters / seeking / clip links (&t=SECONDS) /transcript-with-timestamps (about 1.3x the size of /transcript; x-transcript-cues = line count). Do not summarise a transcript yourself if /summary exists for the video: it is cached and cheaper. CACHING / FRESHNESS / TIMING Transcripts and summaries are cached forever after the first request (header x-cache: HIT | MISS). /semantic is cached too and recomputed only once the background corpus has grown 25% since it was built (x-cache: HIT | MISS, x-corpus-docs = corpus size it was contrasted against). "No captions" is remembered for 6 h (auto-captions can appear hours after upload), then re-checked. ?refresh=1 on any endpoint bypasses the cache for that one request (re-fetches captions / regenerates the summary). Use sparingly: a summary regeneration is an LLM call. First /transcript for a video: ~2-10 s (captions fetched from YouTube). First /summary: ~2 s to first token, 5-30 s total, streamed. Anything cached: < 1 s. Concurrent /summary requests for the same video are coalesced: one generation, everyone else waits and gets the cached result. Batch clients: CPU-heavy steps (yt-dlp extraction, digests) run in a 12-process pool and at most 8 uncached transcripts are fetched from YouTube at once — excess requests queue server-side. ~8-16 parallel requests is the sweet spot; use a client timeout of >= 120 s for uncached videos (a burst of 40 uncached videos takes ~1-2 min end to end); anything cached answers in < 1 s. RESPONSE HEADERS (metadata lives here so bodies stay clean text) all: x-video-id, x-cache (HIT|MISS, transcript & summary), x-transcript-chars transcript: x-transcript-source (yt_dlp_caption_url_rotated | youtube_transcript_api_rotated | seed) summary: x-llm-provider, x-llm-model, x-transcript-truncated (true if the prompt was cut at 500000 chars), x-cache-generated-at semantic: x-semantic-version (dd-v0.2), x-corpus-docs (size of the background corpus the digest was contrasted against) SEMANTIC DIGEST FORMAT (one block per video, ~250 words, designed to be concatenated for a single ranking prompt) VIDEO | TITLE: ... CHANNEL: ... FORM: form guess (monologue/explainer | interview/podcast | tutorial/how-to | listicle/tier | reaction/commentary) | word count | flags ONLY when anomalous: NO-SPEECH, repetitive, promo: , music-tags N TOPICS: distinctive 1-3-word phrases — contrastive tf-idf against every other transcript this service has seen, i.e. what makes THIS video different, not generic vocabulary ENTITIES: proper nouns, products, people NUMBERS: figures with repeat counts (e.g. $300×2) ARC: five equal time slices, each with the terms that peak there — the video's shape over time CENTRAL CLAIMS: up to 7 verbatim sentences the speaker keeps returning to (TextRank + informativeness reweighting), max 2 per fifth of the timeline CATEGORY: YouTube category (Music, Education, Science & Technology, ...) when known — a strong "is this speech?" hint No usable transcript -> /semantic still returns 200 with the VIDEO/TITLE/CHANNEL/CATEGORY header plus "CONTENT: " and header x-transcript: none. (/transcript and /summary return 404 for the same video — they have nothing to give you.) ERRORS 400 bad id / URL 404 no English caption track for this video (/transcript, /summary only) — music videos, no-speech clips, brand-new uploads whose auto-captions aren't generated yet, or small videos YouTube won't caption for logged-out clients; not retryable within hours, nothing to fix 429 YouTube is throttling THIS video's caption download right now (Retry-After: 180) — time-windowed per video/session, not per IP; retry that video after a few minutes, keep going with the rest of your batch 502 upstream failure (YouTube caption fetch or LLM provider, incl. LLM 429s) — retry once after a pause. /summary only answers 200 once the first token exists (it retries the provider first), so a 200 is never empty; a mid-stream failure ends the body with a "[summary interrupted…]" line and is not cached 503 service not configured (rotation credentials / LLM key missing) — report it, don't loop EXAMPLES curl -s https://yt-transcripts.tjwpier.uk/nJMtHxgub2o/semantic curl -s https://yt-transcripts.tjwpier.uk/nJMtHxgub2o/summary curl -s https://yt-transcripts.tjwpier.uk/nJMtHxgub2o/transcript | head -c 4000 curl -sI https://yt-transcripts.tjwpier.uk/nJMtHxgub2o/transcript | grep -i x-transcript-chars # size before loading (HEAD works) curl -s "https://yt-transcripts.tjwpier.uk/?url=https://youtu.be/nJMtHxgub2o" # legacy transcript form curl -s "https://yt-transcripts.tjwpier.uk/nJMtHxgub2o/summary?refresh=1" # force regeneration NOTES Private homelab service (Tom's R730). Transcripts are YouTube caption tracks (creator-provided or auto-generated), not audio transcription — auto captions may lack punctuation and mis-hear names. Summaries are generated by nvidia/nemotron-3-super-120b-a12b:free via openrouter; treat them as a lossy view and verify quotes against /transcript. Transcript views are for research/summarisation; the service does not download video or audio.