feat(podcasts): add YouTube podcast source with transcript-first discovery
New "podcasts" source that discovers podcast content by scanning transcripts from LLM-resolved YouTube channels. Finds content invisible to title-based search — Acquired's "The NFL" episode mentions Taylor Swift 18x, ESPN 117x, Netflix 102x, none in the title. Architecture: - LLM resolves 6-12 podcast channel @handles per topic - Engine fetches recent episodes via yt-dlp (no video download) - Downloads auto-captions and greps for topic keywords - Episodes with 5+ mentions become podcast results with highlights - Runs in parallel, ~15-20s latency, invisible in 3-min research run Pipeline integration: - New source module: scripts/lib/podcast_yt.py - Registered in pipeline, normalizer, signals, planner, render - CLI flag: --podcast-channels=AcquiredFM,lexfridman,... - SOURCE_QUALITY: 0.88 (above YouTube's 0.85) - Opt-in via INCLUDE_SOURCES=podcasts or --search=podcasts Zero new API keys. Zero new dependencies. Reuses yt-dlp + transcript pipeline. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -71,6 +71,7 @@ SOURCE_CAPABILITIES = {
|
||||
"github": {"discussion", "link"},
|
||||
"grounding": {"web", "reference", "link"},
|
||||
"perplexity": {"web", "reference", "analysis"},
|
||||
"podcasts": {"discussion", "video_longform", "expert"},
|
||||
}
|
||||
DEFAULT_INTENT_CAPABILITIES = {
|
||||
"comparison": {"discussion", "video", "web", "reference", "social", "link", "market"},
|
||||
|
||||
Reference in New Issue
Block a user