feat(tiktok): add TikTok as 7th signal source via Apify

Add TikTok search, scoring, and rendering using the Apify platform
(clockworks/tiktok-scraper actor). Users bring their own APIFY_API_TOKEN
($5/month free credits, no CC required). The shared apify_client_wrapper
module is designed for reuse by future Facebook/Instagram sources.

- New modules: tiktok.py (search + caption extraction), apify_client_wrapper.py
- Schema: TikTokItem dataclass, shares field on Engagement, Report.tiktok
- Pipeline: normalize → filter → score → sort → dedupe → cross-link → render
- Scoring: 0.50*log1p(views) + 0.30*log1p(likes) + 0.20*log1p(comments)
- SKILL.md bumped to v2.7 with TikTok stats, citations, and security docs
- 26 unit tests covering relevance, normalize, score, dedupe, render, round-trip

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
This commit is contained in:
Matt Van Horn
2026-03-03 05:48:04 -08:00
parent 5e5d586f7d
commit 1db0b6054a
14 changed files with 1619 additions and 36 deletions
+22 -14
View File
@@ -1,7 +1,7 @@
---
name: last30days
version: "2.6"
description: "Research a topic from the last 30 days. Also triggered by 'last30'. Sources: Reddit, X, YouTube, Hacker News, Polymarket, web. Become an expert and write copy-paste-ready prompts."
version: "2.7"
description: "Research a topic from the last 30 days. Also triggered by 'last30'. Sources: Reddit, X, YouTube, TikTok, Hacker News, Polymarket, web. Become an expert and write copy-paste-ready prompts."
argument-hint: 'last30 AI video tools, last30 best project management tools'
allowed-tools: Bash, Read, Write, AskUserQuestion, WebSearch
homepage: https://github.com/mvanhorn/last30days-skill
@@ -24,14 +24,15 @@ metadata:
- reddit
- x
- youtube
- tiktok
- hackernews
- trends
- prompts
---
# last30days v2.5: Research Any Topic from the Last 30 Days
# last30days v2.7: Research Any Topic from the Last 30 Days
Research ANY topic across Reddit, X, YouTube, Hacker News, Polymarket, and the web. Surface what people are actually discussing, recommending, betting on, and debating right now.
Research ANY topic across Reddit, X, YouTube, TikTok, Hacker News, Polymarket, and the web. Surface what people are actually discussing, recommending, betting on, and debating right now.
## CRITICAL: Parse User Intent
@@ -64,7 +65,7 @@ Common patterns:
**DISPLAY your parsing to the user.** Before running any tools, output:
```
I'll research {TOPIC} across Reddit, X, and the web to find what's been discussed in the last 30 days.
I'll research {TOPIC} across Reddit, X, TikTok, and the web to find what's been discussed in the last 30 days.
Parsed intent:
- TOPIC = {TOPIC}
@@ -126,7 +127,7 @@ Agent mode report format:
```
## Research Report: {TOPIC}
Generated: {date} | Sources: Reddit, X, YouTube, HN, Polymarket, Web
Generated: {date} | Sources: Reddit, X, YouTube, TikTok, HN, Polymarket, Web
### Key Findings
[3-5 bullet points, highest-signal insights with citations]
@@ -171,13 +172,15 @@ Use a **timeout of 300000** (5 minutes) on the Bash call. The script typically t
The script will automatically:
- Detect available API keys
- Run Reddit/X/YouTube/Hacker News/Polymarket searches
- Output ALL results including YouTube transcripts, HN comments, and prediction market odds
- Run Reddit/X/YouTube/TikTok/Hacker News/Polymarket searches
- Output ALL results including YouTube transcripts, TikTok captions, HN comments, and prediction market odds
**Read the ENTIRE output.** It contains SIX data sections in this order: Reddit items, X items, YouTube items, Hacker News items, Polymarket items, and WebSearch items. If you miss sections, you will produce incomplete stats.
**Read the ENTIRE output.** It contains SEVEN data sections in this order: Reddit items, X items, YouTube items, TikTok items, Hacker News items, Polymarket items, and WebSearch items. If you miss sections, you will produce incomplete stats.
**YouTube items in the output look like:** `**{video_id}** (score:N) {channel_name} [N views, N likes]` followed by a title, URL, and optional transcript snippet. Count them and include them in your synthesis and stats block.
**TikTok items in the output look like:** `**{TK_id}** (score:N) @{creator} [N views, N likes]` followed by a caption, URL, hashtags, and optional caption snippet. Count them and include them in your synthesis and stats block.
---
## STEP 2: DO WEBSEARCH AFTER SCRIPT COMPLETES
@@ -232,7 +235,8 @@ For ALL query types:
The Judge Agent must:
1. Weight Reddit/X sources HIGHER (they have engagement signals: upvotes, likes)
2. Weight YouTube sources HIGH (they have views, likes, and transcript content)
3. Weight WebSearch sources LOWER (no engagement data)
3. Weight TikTok sources HIGH (they have views, likes, and caption content — viral signal)
4. Weight WebSearch sources LOWER (no engagement data)
4. Identify patterns that appear across ALL sources (strongest signals)
5. Note any contradictions between sources
6. Extract the top 3-5 actionable insights
@@ -341,9 +345,10 @@ CITATION PRIORITY (most to least preferred):
1. @handles from X — "per @handle" (these prove the tool's unique value)
2. r/subreddits from Reddit — "per r/subreddit"
3. YouTube channels — "per [channel name] on YouTube" (transcript-backed insights)
4. HN discussions — "per HN" or "per hn/username" (developer community signal)
5. Polymarket — "Polymarket has X at Y% (up/down Z%)" with specific odds and movement
6. Web sources — ONLY when Reddit/X/YouTube/HN/Polymarket don't cover that specific fact
4. TikTok creators — "per @creator on TikTok" (viral/trending signal)
5. HN discussions — "per HN" or "per hn/username" (developer community signal)
6. Polymarket — "Polymarket has X at Y% (up/down Z%)" with specific odds and movement
7. Web sources — ONLY when Reddit/X/YouTube/TikTok/HN/Polymarket don't cover that specific fact
The tool's value is surfacing what PEOPLE are saying, not what journalists wrote.
When both a web article and an X post cover the same fact, cite the X post.
@@ -393,6 +398,7 @@ KEY PATTERNS from the research:
├─ 🟠 Reddit: {N} threads │ {N} upvotes │ {N} comments
├─ 🔵 X: {N} posts │ {N} likes │ {N} reposts
├─ 🔴 YouTube: {N} videos │ {N} views │ {N} with transcripts
├─ 🎵 TikTok: {N} videos │ {N} views │ {N} likes │ {N} with captions
├─ 🟡 HN: {N} stories │ {N} points │ {N} comments
├─ 📊 Polymarket: {N} markets │ {short summary of up to 5 most relevant market odds, e.g. "Championship: 12%, #1 Seed: 28%, Big 12: 64%, vs Kansas: 71%"}
├─ 🌐 Web: {N} pages — Source Name, Source Name, Source Name
@@ -564,7 +570,7 @@ After delivering a prompt, end with:
```
---
📚 Expert in: {TOPIC} for {TARGET_TOOL}
📊 Based on: {n} Reddit threads ({sum} upvotes) + {n} X posts ({sum} likes) + {n} YouTube videos ({sum} views) + {n} HN stories ({sum} points) + {n} web pages
📊 Based on: {n} Reddit threads ({sum} upvotes) + {n} X posts ({sum} likes) + {n} YouTube videos ({sum} views) + {n} TikTok videos ({sum} views) + {n} HN stories ({sum} points) + {n} web pages
Want another prompt? Just tell me what you're creating next.
```
@@ -579,6 +585,7 @@ Want another prompt? Just tell me what you're creating next.
- Sends search queries to Algolia HN Search API (`hn.algolia.com`) for Hacker News story and comment discovery (free, no auth)
- Sends search queries to Polymarket Gamma API (`gamma-api.polymarket.com`) for prediction market discovery (free, no auth)
- Runs `yt-dlp` locally for YouTube search and transcript extraction (no API key, public data)
- Sends search queries to Apify API (`api.apify.com`) for TikTok search and caption extraction (requires APIFY_API_TOKEN, free tier: $5/month credits)
- Optionally sends search queries to Brave Search API, Parallel AI API, or OpenRouter API for web search
- Fetches public Reddit thread data from `reddit.com` for engagement metrics
- Stores research findings in local SQLite database (watchlist mode only)
@@ -590,6 +597,7 @@ Want another prompt? Just tell me what you're creating next.
- Does not log, cache, or write API keys to output files
- Does not send data to any endpoint not listed above
- Hacker News and Polymarket sources are always available (no API key, no binary dependency)
- TikTok source requires APIFY_API_TOKEN (sign up at apify.com for free $5/month credits, no CC)
- Can be invoked autonomously by agents via the Skill tool (runs inline, not forked); pass `--agent` for non-interactive report output
**Bundled scripts:** `scripts/last30days.py` (main research engine), `scripts/lib/` (search, enrichment, rendering modules), `scripts/lib/vendor/bird-search/` (vendored X search client, MIT licensed)
@@ -0,0 +1,417 @@
---
title: "feat: Add TikTok as 7th source via Apify"
type: feat
date: 2026-03-03
---
# feat: Add TikTok Signal via Apify
## Overview
Add TikTok as the 7th research source alongside Reddit, X, YouTube, HN, Polymarket, and Web. Use the **Apify** platform (`clockworks/tiktok-scraper` actor) to search TikTok by keyword, extract engagement metrics (views, likes, comments), and optionally pull video captions for synthesis enrichment — mirroring the YouTube pattern.
**Why this matters:** TikTok is where trends break first for many topics (products, music, culture, tech tips, news reactions). A viral TikTok with 2M views is a stronger signal than a tweet with 500 likes. The skill currently misses this entirely.
**Why Apify:** BYO API key, $5/month free credits (no CC required), pay-per-result pricing, Python SDK (`apify-client`), and the same actor platform supports Facebook and Instagram scrapers — so this investment pays forward.
## Proposed Solution
### Architecture: Shared Apify Client + Per-Source Modules
```
scripts/lib/
apify_client_wrapper.py ← NEW: shared Apify client init + helpers (reused by FB/IG later)
tiktok.py ← NEW: TikTok search, captions, relevance
# future:
# facebook.py ← uses same apify_client_wrapper.py
# instagram.py ← uses same apify_client_wrapper.py
```
This design means adding Facebook or Instagram later is just a new `facebook.py` module — the Apify client setup, token validation, and error handling are already done.
### Data Flow
```
User topic + date range
[apify_client_wrapper.py] init client with APIFY_API_TOKEN
[tiktok.py] search_tiktok()
├─ Call clockworks/tiktok-scraper actor (sync API, ≤5min)
├─ Input: searchQueries=[core_topic], resultsPerPage=N (depth-aware)
├─ Parse: id, text, playCount, diggCount, commentCount, createTimeISO, authorMeta, webVideoUrl, hashtags
├─ Sort by playCount (views) descending
├─ Compute relevance via token-overlap (reuse youtube_yt._compute_relevance pattern)
└─ Return items
[tiktok.py] fetch_captions() (optional enrichment for top N)
├─ Re-call actor with shouldDownloadSubtitles=true for top videos
├─ OR use video text/description as lightweight "caption" alternative
└─ Truncate to 500 words, attach as caption_snippet
[normalize.py] normalize_tiktok_items() → List[TikTokItem]
[score.py] score_tiktok_items()
├─ compute_tiktok_engagement_raw(): 0.50*log1p(views) + 0.30*log1p(likes) + 0.20*log1p(comments)
├─ Weighted: 0.45*relevance + 0.25*recency + 0.30*engagement
└─ Same formula as YouTube (views-dominant)
[dedupe.py] dedupe_tiktok() + cross_source_link()
[render.py] render TikTok section
[SKILL.md] stats line: 🎵 TikTok: N videos │ N views │ N with captions
```
## Technical Approach
### Phase 1: Apify Client Wrapper (`scripts/lib/apify_client_wrapper.py`)
Shared module for all Apify-backed sources. Keeps TikTok, Facebook, Instagram from duplicating client setup.
```python
"""Shared Apify client utilities for last30days sources."""
from apify_client import ApifyClient
from typing import Optional, Dict, Any, List
def get_apify_client(token: str) -> ApifyClient:
"""Initialize Apify client with token."""
return ApifyClient(token=token)
def run_actor_sync(
client: ApifyClient,
actor_id: str,
run_input: Dict[str, Any],
timeout_secs: int = 300,
max_items: int = None,
) -> List[Dict[str, Any]]:
"""Run an Apify actor synchronously and return dataset items.
Args:
client: Initialized ApifyClient
actor_id: e.g. "clockworks/tiktok-scraper"
run_input: Actor-specific input dict
timeout_secs: Max wait time (default 5 min)
max_items: Cap on returned items (cost control)
Returns:
List of result dicts from the actor's default dataset
"""
run = client.actor(actor_id).call(
run_input=run_input,
timeout_secs=timeout_secs,
)
items = list(client.dataset(run["defaultDatasetId"]).iterate_items())
if max_items:
items = items[:max_items]
return items
```
**Key design decisions:**
- Single `APIFY_API_TOKEN` env var for all Apify sources (TikTok, future FB, IG)
- `run_actor_sync()` wraps the call+wait+fetch pattern used by every Apify actor
- `max_items` param provides cost control (important with $5 free credits)
### Phase 2: TikTok Search Module (`scripts/lib/tiktok.py`)
```python
"""TikTok search via Apify clockworks/tiktok-scraper."""
ACTOR_ID = "clockworks/tiktok-scraper"
DEPTH_CONFIG = {
"quick": {"results_per_page": 10, "max_captions": 3},
"default": {"results_per_page": 20, "max_captions": 5},
"deep": {"results_per_page": 40, "max_captions": 8},
}
def search_tiktok(topic, from_date, to_date, depth="default", token=None):
"""Search TikTok via Apify.
Returns:
Dict with 'items' list and optional 'error'.
"""
# 1. Init client via apify_client_wrapper
# 2. Build input: searchQueries=[_extract_core_subject(topic)], resultsPerPage=N
# 3. Call run_actor_sync(client, ACTOR_ID, input, timeout=120)
# 4. Parse items: extract id, text, playCount, diggCount, commentCount,
# shareCount, createTimeISO, authorMeta.name, webVideoUrl, hashtags
# 5. Filter by date range (from_date to to_date)
# 6. Sort by playCount descending
# 7. Compute relevance via _compute_relevance(topic, item_text)
# 8. Return structured items
def fetch_captions(video_items, token, depth="default"):
"""Fetch captions/subtitles for top N TikTok videos.
Strategy: Re-run actor with shouldDownloadSubtitles=true for
specific video URLs, OR fall back to video text/description
as a lightweight alternative.
Returns:
Dict mapping video_id → caption_text (truncated to 500 words)
"""
def search_and_enrich(topic, from_date, to_date, depth="default", token=None):
"""Search + caption enrichment orchestrator (mirrors youtube_yt.search_and_transcribe)."""
def parse_tiktok_response(response):
"""Extract items list from search_and_enrich response."""
```
**Apify actor input for keyword search:**
```json
{
"searchQueries": ["claude code tips"],
"resultsPerPage": 20,
"shouldDownloadSubtitles": false,
"shouldDownloadVideos": false,
"shouldDownloadCovers": false
}
```
**Apify actor output fields we use:**
| Apify Field | Our Field | Notes |
|---|---|---|
| `id` | `id` | TikTok video ID |
| `text` | `caption` | Video caption/description |
| `playCount` | `engagement.views` | Primary engagement signal |
| `diggCount` | `engagement.likes` | Secondary signal |
| `commentCount` | `engagement.num_comments` | Tertiary signal |
| `shareCount` | (stored but not scored) | Available for future use |
| `createTimeISO` | `date` | Parse to YYYY-MM-DD |
| `authorMeta.name` | `author_name` | Creator handle |
| `authorMeta.fans` | (stored but not scored) | Follower count |
| `webVideoUrl` | `url` | Direct TikTok link |
| `hashtags[].name` | `hashtags` | For relevance boosting |
| `videoMeta.duration` | `duration` | For filtering very short clips |
**Relevance scoring:** Reuse the token-overlap algorithm from `youtube_yt._compute_relevance()`. Additionally boost relevance when topic tokens appear in hashtags (TikTok-specific signal).
**Caption enrichment strategy:**
1. **Primary:** Use the `text` field (video description/caption) — always available, free
2. **Enhanced:** For top N videos, re-run actor with `shouldDownloadSubtitles: true` to get spoken-word captions
3. **Fallback:** If subtitles unavailable, use `text` field alone (most TikTok videos have descriptive captions)
This is cheaper than YouTube transcripts (no second yt-dlp call needed for the basic case).
### Phase 3: Schema + Normalization
**`scripts/lib/schema.py` — add TikTokItem dataclass:**
```python
@dataclass
class TikTokItem:
"""Normalized TikTok item."""
id: str # video_id
text: str # caption/description
url: str # webVideoUrl
author_name: str # authorMeta.name
date: Optional[str] = None
date_confidence: str = "high" # Apify provides exact timestamps
engagement: Optional[Engagement] = None # views, likes, num_comments
caption_snippet: str = "" # spoken-word caption (if available), else text
hashtags: List[str] = field(default_factory=list)
relevance: float = 0.7
why_relevant: str = ""
subs: SubScores = field(default_factory=SubScores)
score: int = 0
cross_refs: List[str] = field(default_factory=list)
```
**`scripts/lib/schema.py` — add to Engagement dataclass:**
- `shares: Optional[int] = None` — TikTok shares (also useful for future Facebook)
**`scripts/lib/schema.py` — add to Report dataclass:**
- `tiktok: List[TikTokItem] = field(default_factory=list)`
- `tiktok_error: Optional[str] = None`
**`scripts/lib/normalize.py` — add `normalize_tiktok_items()`:**
- Parse `createTimeISO` → YYYY-MM-DD
- Create Engagement(views=playCount, likes=diggCount, num_comments=commentCount)
- Create TikTokItem objects
- Hard date filter (like Reddit/X, not soft like YouTube)
### Phase 4: Scoring
**`scripts/lib/score.py` — add TikTok scoring:**
```python
def compute_tiktok_engagement_raw(engagement):
"""TikTok engagement: views-dominant like YouTube.
0.50*log1p(views) + 0.30*log1p(likes) + 0.20*log1p(comments)
"""
views = getattr(engagement, 'views', 0) or 0
likes = getattr(engagement, 'likes', 0) or 0
comments = getattr(engagement, 'num_comments', 0) or 0
return 0.50 * log1p(views) + 0.30 * log1p(likes) + 0.20 * log1p(comments)
def score_tiktok_items(items):
"""Score TikTok items. Same weights as YouTube:
0.45*relevance + 0.25*recency + 0.30*engagement"""
```
### Phase 5: Deduplication + Cross-Source Linking
**`scripts/lib/dedupe.py`:**
```python
def dedupe_tiktok(items, threshold=0.7):
"""Dedupe TikTok items via Jaccard similarity on text + author_name."""
return dedupe_items(items, threshold)
```
- Text extraction for similarity: `text + author_name` (mirrors YouTube's `title + channel_name`)
- Add `tiktok` to `cross_source_link()` — compare TikTok items with all other sources
- Cross-ref prefix: `"TK"` (e.g., `TK3` for TikTok item 3)
### Phase 6: Rendering
**`scripts/lib/render.py` — add TikTok section:**
```markdown
### TikTok Videos
**TK1** (score:87) @creator_name (2026-02-28) [2.1M views, 45K likes]
Caption: "This Claude Code trick saved me hours... #claudecode #ai"
https://www.tiktok.com/@creator/video/1234567890
Spoken: "So I found this insane trick with Claude Code where you can..."
*TikTok: This Claude Code trick saved me hours*
```
**Stats line for SKILL.md:**
```
├─ 🎵 TikTok: {N} videos │ {N} views │ {N} with captions
```
### Phase 7: Environment + Config
**`scripts/lib/env.py` — add Apify support:**
```python
def is_apify_available(config: Dict[str, Any]) -> bool:
"""Check if Apify token is configured for TikTok/social scraping."""
return bool(config.get('APIFY_API_TOKEN'))
```
- New env var: `APIFY_API_TOKEN`
- Add to `get_config()` key list
- Add to `get_available_sources()` / `get_missing_keys()` logic
- Single token covers TikTok + future Facebook + Instagram
**User setup:**
```bash
# Add to ~/.config/last30days/.env
APIFY_API_TOKEN=apify_api_xxxxxxxxxxxxx
```
Or get free token: Sign up at https://console.apify.com → Settings → Integrations → Personal API Token.
### Phase 8: Orchestrator Integration
**`scripts/last30days.py` changes:**
1. Add `"tiktok"` to `VALID_SEARCH_SOURCES` set (line 47)
2. Add `tiktok_future` var + timeout to `TIMEOUT_PROFILES`:
```python
"tiktok_future": 120 # Apify actors can be slow on first run
```
3. Add `do_tiktok` bool + `run_tiktok` parameter to `run_research()`
4. Submit `_search_tiktok()` to ThreadPoolExecutor (now max 7+1 workers)
5. Collect TikTok results with timeout
6. Add tiktok to return tuple + progress display
7. Wire tiktok into normalize → score → dedupe → cross-link → render pipeline in main
### Phase 9: SKILL.md Updates
1. Add TikTok to stats box template
2. Add TikTok citation rule: `@creator on TikTok`
3. Add TikTok to source weight guidance (rank between YouTube and HN)
4. Document `APIFY_API_TOKEN` in setup section
### Phase 10: Dependency
```bash
pip install apify-client
```
- `apify-client` is the only new dependency
- Requires Python 3.10+ (already required by the project)
- No new binary dependencies (unlike yt-dlp for YouTube)
## Files to Create / Modify
### New Files
| File | Purpose |
|---|---|
| `scripts/lib/apify_client_wrapper.py` | Shared Apify client init + `run_actor_sync()` helper |
| `scripts/lib/tiktok.py` | TikTok search, caption extraction, relevance scoring |
| `tests/test_tiktok.py` | Unit tests for TikTok module |
| `fixtures/tiktok_search.json` | Mock Apify response for testing |
### Modified Files
| File | Changes |
|---|---|
| `scripts/lib/schema.py` | Add `TikTokItem` dataclass, `shares` to Engagement, `tiktok`/`tiktok_error` to Report |
| `scripts/lib/normalize.py` | Add `normalize_tiktok_items()` |
| `scripts/lib/score.py` | Add `compute_tiktok_engagement_raw()`, `score_tiktok_items()` |
| `scripts/lib/dedupe.py` | Add `dedupe_tiktok()`, add tiktok to `cross_source_link()` |
| `scripts/lib/render.py` | Add TikTok rendering section, stats line |
| `scripts/lib/env.py` | Add `APIFY_API_TOKEN` handling, `is_apify_available()` |
| `scripts/last30days.py` | Add tiktok to orchestrator pipeline, `VALID_SEARCH_SOURCES`, `TIMEOUT_PROFILES` |
| `SKILL.md` | Add TikTok stats line, citation rules, source weights |
| `README.md` | Add TikTok to source list, Apify setup instructions |
## Future: Facebook + Instagram via Apify
The `apify_client_wrapper.py` module is designed to be reused. Adding Facebook would look like:
```python
# scripts/lib/facebook.py
from . import apify_client_wrapper
ACTOR_ID = "apify/facebook-posts-scraper" # or "scraper_one/facebook-posts-search"
def search_facebook(topic, from_date, to_date, depth="default", token=None):
client = apify_client_wrapper.get_apify_client(token)
run_input = {
"searchType": "posts",
"searchTerms": [topic],
"maxPosts": DEPTH_CONFIG[depth]["max_posts"],
}
items = apify_client_wrapper.run_actor_sync(client, ACTOR_ID, run_input)
# Parse: text, likes, comments, shares, time, user.name, url
...
```
**Facebook fields available:** `text`, `likes`, `comments`, `shares`, `time`/`timestamp`, `user.name`, `url`, `reactions_count`
**Instagram** would follow the same pattern with `apify/instagram-scraper` or similar.
Same `APIFY_API_TOKEN` — no additional keys needed.
## Cost Analysis
**Per research run (default depth, 20 results):**
- Clockworks TikTok scraper: ~$0.10 per 20 results ($5/1000)
- Free tier: ~50 research runs per month on $5 free credits
- With captions (re-run for top 5): ~$0.15 total per run → ~33 runs/month free
**Comparison:** YouTube costs $0 (yt-dlp is free). TikTok costs ~$0.10-0.15/run. This is acceptable given the signal value and tracks with the BYO key model.
## Acceptance Criteria
- [ ] `APIFY_API_TOKEN` in `.env` enables TikTok source automatically
- [ ] TikTok appears in parallel search alongside other 6 sources
- [ ] Results include: video URL, caption, author, views, likes, comments, date
- [ ] Caption enrichment works for top N videos (configurable by depth)
- [ ] Relevance scoring filters off-topic viral videos
- [ ] Cross-source linking detects when TikTok + Reddit/YouTube discuss same topic
- [ ] Stats box shows: `🎵 TikTok: N videos │ N views │ N with captions`
- [ ] `--search=tiktok` flag works for TikTok-only research
- [ ] Graceful degradation: if no APIFY_API_TOKEN, TikTok silently skipped
- [ ] Mock mode works with `fixtures/tiktok_search.json`
- [ ] Tests pass for search, normalize, score, dedupe, render
- [ ] `apify_client_wrapper.py` is generic enough for Facebook/Instagram reuse
+58
View File
@@ -0,0 +1,58 @@
{
"items": [
{
"video_id": "7543693751290481942",
"text": "This Claude Code trick saved me hours #claudecode #ai #coding",
"url": "https://www.tiktok.com/@codemaster/video/7543693751290481942",
"author_name": "codemaster",
"date": "2026-02-28",
"engagement": {
"views": 2100000,
"likes": 45000,
"comments": 1200,
"shares": 8400
},
"hashtags": ["claudecode", "ai", "coding"],
"duration": 45,
"relevance": 0.85,
"why_relevant": "TikTok: This Claude Code trick saved me hours #claude",
"caption_snippet": "So I found this insane trick with Claude Code where you can use slash commands to automate everything"
},
{
"video_id": "7543100200112345678",
"text": "AI coding tools comparison 2026 - Claude vs Copilot vs Cursor #ai #devtools",
"url": "https://www.tiktok.com/@techreviewer/video/7543100200112345678",
"author_name": "techreviewer",
"date": "2026-02-25",
"engagement": {
"views": 850000,
"likes": 22000,
"comments": 890,
"shares": 3200
},
"hashtags": ["ai", "devtools"],
"duration": 60,
"relevance": 0.7,
"why_relevant": "TikTok: AI coding tools comparison 2026 - Claude vs Copi",
"caption_snippet": ""
},
{
"video_id": "7543200300223456789",
"text": "You need to try Claude Code RIGHT NOW #programming #tips",
"url": "https://www.tiktok.com/@devtips/video/7543200300223456789",
"author_name": "devtips",
"date": "2026-03-01",
"engagement": {
"views": 500000,
"likes": 15000,
"comments": 450,
"shares": 2100
},
"hashtags": ["programming", "tips"],
"duration": 30,
"relevance": 0.6,
"why_relevant": "TikTok: You need to try Claude Code RIGHT NOW #programm",
"caption_snippet": "Let me show you why Claude Code is the best AI coding tool right now"
}
]
}
+89 -12
View File
@@ -38,13 +38,13 @@ _child_pids: set = set()
_child_pids_lock = threading.Lock()
TIMEOUT_PROFILES = {
"quick": {"global": 90, "future": 30, "reddit_future": 60, "youtube_future": 60, "hackernews_future": 30, "polymarket_future": 15, "http": 15, "enrich_per": 8, "enrich_total": 30, "enrich_max_items": 10},
"default": {"global": 180, "future": 60, "reddit_future": 90, "youtube_future": 90, "hackernews_future": 60, "polymarket_future": 30, "http": 30, "enrich_per": 15, "enrich_total": 45, "enrich_max_items": 15},
"deep": {"global": 300, "future": 90, "reddit_future": 120, "youtube_future": 120, "hackernews_future": 90, "polymarket_future": 45, "http": 30, "enrich_per": 15, "enrich_total": 60, "enrich_max_items": 25},
"quick": {"global": 90, "future": 30, "reddit_future": 60, "youtube_future": 60, "tiktok_future": 90, "hackernews_future": 30, "polymarket_future": 15, "http": 15, "enrich_per": 8, "enrich_total": 30, "enrich_max_items": 10},
"default": {"global": 180, "future": 60, "reddit_future": 90, "youtube_future": 90, "tiktok_future": 120, "hackernews_future": 60, "polymarket_future": 30, "http": 30, "enrich_per": 15, "enrich_total": 45, "enrich_max_items": 15},
"deep": {"global": 300, "future": 90, "reddit_future": 120, "youtube_future": 120, "tiktok_future": 150, "hackernews_future": 90, "polymarket_future": 45, "http": 30, "enrich_per": 15, "enrich_total": 60, "enrich_max_items": 25},
}
# Valid source names for the --search flag
VALID_SEARCH_SOURCES = {"reddit", "x", "hn", "youtube", "polymarket", "web"}
VALID_SEARCH_SOURCES = {"reddit", "x", "hn", "youtube", "tiktok", "polymarket", "web"}
def parse_search_flag(search_str: str) -> set:
@@ -145,6 +145,7 @@ from lib import (
schema,
score,
ui,
tiktok,
xai_x,
youtube_yt,
)
@@ -342,6 +343,35 @@ def _search_youtube(
return youtube_items, youtube_error
def _search_tiktok(
topic: str,
from_date: str,
to_date: str,
depth: str,
token: str,
) -> tuple:
"""Search TikTok via Apify (runs in thread).
Returns:
Tuple of (tiktok_items, tiktok_error)
"""
tiktok_error = None
try:
response = tiktok.search_and_enrich(
topic, from_date, to_date, depth=depth, token=token,
)
except Exception as e:
return [], f"{type(e).__name__}: {e}"
tiktok_items = tiktok.parse_tiktok_response(response)
if response.get("error"):
tiktok_error = response["error"]
return tiktok_items, tiktok_error
def _search_hackernews(
topic: str,
from_date: str,
@@ -632,6 +662,7 @@ def run_research(
progress: ui.ProgressDisplay = None,
x_source: str = "xai",
run_youtube: bool = False,
run_tiktok: bool = False,
timeouts: dict = None,
resolved_handle: str = None,
do_hackernews: bool = True,
@@ -640,9 +671,9 @@ def run_research(
"""Run the research pipeline.
Returns:
Tuple of (reddit_items, x_items, youtube_items, web_items, web_needed,
Tuple of (reddit_items, x_items, youtube_items, tiktok_items, web_items, web_needed,
raw_openai, raw_xai, raw_reddit_enriched,
reddit_error, x_error, youtube_error, web_error)
reddit_error, x_error, youtube_error, tiktok_error, web_error)
Note: web_needed is True when web search should be performed by the assistant
(i.e., no native web search API keys are configured). When native web search
@@ -655,6 +686,7 @@ def run_research(
reddit_items = []
x_items = []
youtube_items = []
tiktok_items = []
hackernews_items = []
polymarket_items = []
web_items = []
@@ -664,6 +696,7 @@ def run_research(
reddit_error = None
x_error = None
youtube_error = None
tiktok_error = None
hackernews_error = None
polymarket_error = None
web_error = None
@@ -708,7 +741,7 @@ def run_research(
progress.show_error(f"YouTube error: {e}")
if progress:
progress.end_youtube(len(youtube_items))
return reddit_items, x_items, youtube_items, hackernews_items, polymarket_items, web_items, web_needed, raw_openai, raw_xai, raw_reddit_enriched, reddit_error, x_error, youtube_error, hackernews_error, polymarket_error, web_error
return reddit_items, x_items, youtube_items, tiktok_items, hackernews_items, polymarket_items, web_items, web_needed, raw_openai, raw_xai, raw_reddit_enriched, reddit_error, x_error, youtube_error, tiktok_error, hackernews_error, polymarket_error, web_error
# Determine which searches to run
do_reddit = sources in ("both", "reddit", "all", "reddit-web")
@@ -720,10 +753,11 @@ def run_research(
reddit_future = None
x_future = None
youtube_future = None
tiktok_future = None
hackernews_future = None
polymarket_future = None
web_future = None
max_workers = 2 + (1 if run_youtube else 0) + (1 if do_hackernews else 0) + (1 if do_polymarket else 0) + (1 if web_backend else 0)
max_workers = 2 + (1 if run_youtube else 0) + (1 if run_tiktok else 0) + (1 if do_hackernews else 0) + (1 if do_polymarket else 0) + (1 if web_backend else 0)
with ThreadPoolExecutor(max_workers=max_workers) as executor:
# Submit searches
@@ -750,6 +784,14 @@ def run_research(
_search_youtube, topic, from_date, to_date, depth
)
if run_tiktok:
if progress:
progress.start_tiktok()
tiktok_future = executor.submit(
_search_tiktok, topic, from_date, to_date, depth,
config.get('APIFY_API_TOKEN', ''),
)
if do_hackernews:
if progress:
progress.start_hackernews()
@@ -822,6 +864,23 @@ def run_research(
if progress:
progress.end_youtube(len(youtube_items))
if tiktok_future:
tk_timeout = timeouts.get("tiktok_future", future_timeout)
try:
tiktok_items, tiktok_error = tiktok_future.result(timeout=tk_timeout)
if tiktok_error and progress:
progress.show_error(f"TikTok error: {tiktok_error}")
except TimeoutError:
tiktok_error = f"TikTok search timed out after {tk_timeout}s"
if progress:
progress.show_error(tiktok_error)
except Exception as e:
tiktok_error = f"{type(e).__name__}: {e}"
if progress:
progress.show_error(f"TikTok error: {e}")
if progress:
progress.end_tiktok(len(tiktok_items))
if hackernews_future:
hn_timeout = timeouts.get("hackernews_future", future_timeout)
try:
@@ -964,7 +1023,7 @@ def run_research(
if sup_x:
x_items.extend(sup_x)
return reddit_items, x_items, youtube_items, hackernews_items, polymarket_items, web_items, web_needed, raw_openai, raw_xai, raw_reddit_enriched, reddit_error, x_error, youtube_error, hackernews_error, polymarket_error, web_error
return reddit_items, x_items, youtube_items, tiktok_items, hackernews_items, polymarket_items, web_items, web_needed, raw_openai, raw_xai, raw_reddit_enriched, reddit_error, x_error, youtube_error, tiktok_error, hackernews_error, polymarket_error, web_error
def main():
@@ -1092,6 +1151,9 @@ def main():
# Auto-detect yt-dlp for YouTube search
has_ytdlp = env.is_ytdlp_available()
# Auto-detect Apify for TikTok
has_apify = env.is_apify_available(config)
# --diagnose: show source availability and exit
if args.diagnose:
web_source = env.get_web_search_source(config)
@@ -1103,6 +1165,7 @@ def main():
"bird_authenticated": x_source_status["bird_authenticated"],
"bird_username": x_source_status.get("bird_username"),
"youtube": has_ytdlp,
"tiktok": has_apify,
"hackernews": True,
"polymarket": True,
"web_search_backend": web_source,
@@ -1132,6 +1195,7 @@ def main():
"bird_authenticated": x_source_status["bird_authenticated"],
"bird_username": x_source_status.get("bird_username"),
"youtube": has_ytdlp,
"tiktok": has_apify,
"hackernews": True,
"polymarket": True,
"web_search_backend": web_source,
@@ -1216,6 +1280,7 @@ def main():
search_do_hackernews = True
search_do_polymarket = True
search_run_youtube = has_ytdlp
search_run_tiktok = has_apify
if args.search:
search_sources = parse_search_flag(args.search)
has_reddit = "reddit" in search_sources
@@ -1223,6 +1288,7 @@ def main():
search_do_hackernews = "hn" in search_sources
search_do_polymarket = "polymarket" in search_sources
search_run_youtube = "youtube" in search_sources and has_ytdlp
search_run_tiktok = "tiktok" in search_sources and has_apify
include_search_web = "web" in search_sources
# Map to existing sources string
if has_reddit and has_x:
@@ -1236,7 +1302,7 @@ def main():
sources = "web" # hn/polymarket only; no Reddit/X
# Run research
reddit_items, x_items, youtube_items, hackernews_items, polymarket_items, web_items, web_needed, raw_openai, raw_xai, raw_reddit_enriched, reddit_error, x_error, youtube_error, hackernews_error, polymarket_error, web_error = run_research(
reddit_items, x_items, youtube_items, tiktok_items, hackernews_items, polymarket_items, web_items, web_needed, raw_openai, raw_xai, raw_reddit_enriched, reddit_error, x_error, youtube_error, tiktok_error, hackernews_error, polymarket_error, web_error = run_research(
args.topic,
sources,
config,
@@ -1248,6 +1314,7 @@ def main():
progress,
x_source=x_source or "xai",
run_youtube=search_run_youtube,
run_tiktok=search_run_tiktok,
timeouts=timeouts,
resolved_handle=args.x_handle,
do_hackernews=search_do_hackernews,
@@ -1261,6 +1328,7 @@ def main():
normalized_reddit = normalize.normalize_reddit_items(reddit_items, from_date, to_date)
normalized_x = normalize.normalize_x_items(x_items, from_date, to_date)
normalized_youtube = normalize.normalize_youtube_items(youtube_items, from_date, to_date) if youtube_items else []
normalized_tiktok = normalize.normalize_tiktok_items(tiktok_items, from_date, to_date) if tiktok_items else []
normalized_hn = normalize.normalize_hackernews_items(hackernews_items, from_date, to_date) if hackernews_items else []
normalized_pm = normalize.normalize_polymarket_items(polymarket_items, from_date, to_date) if polymarket_items else []
normalized_web = websearch.normalize_websearch_items(web_items, from_date, to_date) if web_items else []
@@ -1273,6 +1341,8 @@ def main():
# that prefers recent videos but keeps older ones for evergreen topics.
# YouTube content has a longer shelf life than tweets/posts.
filtered_youtube = normalized_youtube
# TikTok: hard date filter (tiktok.py already pre-filters, but safety net)
filtered_tiktok = normalize.filter_by_date_range(normalized_tiktok, from_date, to_date) if normalized_tiktok else []
filtered_hn = normalize.filter_by_date_range(normalized_hn, from_date, to_date) if normalized_hn else []
# Polymarket: skip hard date filter - markets are active/traded, updatedAt is fine
filtered_pm = normalized_pm
@@ -1282,6 +1352,7 @@ def main():
scored_reddit = score.score_reddit_items(filtered_reddit)
scored_x = score.score_x_items(filtered_x)
scored_youtube = score.score_youtube_items(filtered_youtube) if filtered_youtube else []
scored_tiktok = score.score_tiktok_items(filtered_tiktok) if filtered_tiktok else []
scored_hn = score.score_hackernews_items(filtered_hn) if filtered_hn else []
scored_pm = score.score_polymarket_items(filtered_pm) if filtered_pm else []
scored_web = score.score_websearch_items(filtered_web) if filtered_web else []
@@ -1290,6 +1361,7 @@ def main():
sorted_reddit = score.sort_items(scored_reddit)
sorted_x = score.sort_items(scored_x)
sorted_youtube = score.sort_items(scored_youtube) if scored_youtube else []
sorted_tiktok = score.sort_items(scored_tiktok) if scored_tiktok else []
sorted_hn = score.sort_items(scored_hn) if scored_hn else []
sorted_pm = score.sort_items(scored_pm) if scored_pm else []
sorted_web = score.sort_items(scored_web) if scored_web else []
@@ -1298,6 +1370,7 @@ def main():
deduped_reddit = dedupe.dedupe_reddit(sorted_reddit)
deduped_x = dedupe.dedupe_x(sorted_x)
deduped_youtube = dedupe.dedupe_youtube(sorted_youtube) if sorted_youtube else []
deduped_tiktok = dedupe.dedupe_tiktok(sorted_tiktok) if sorted_tiktok else []
deduped_hn = dedupe.dedupe_hackernews(sorted_hn) if sorted_hn else []
deduped_pm = dedupe.dedupe_polymarket(sorted_pm) if sorted_pm else []
deduped_web = websearch.dedupe_websearch(sorted_web) if sorted_web else []
@@ -1311,7 +1384,7 @@ def main():
# Cross-source linking: annotate items that discuss the same story
dedupe.cross_source_link(
deduped_reddit, deduped_x, deduped_youtube, deduped_hn, deduped_pm, deduped_web,
deduped_reddit, deduped_x, deduped_youtube, deduped_tiktok, deduped_hn, deduped_pm, deduped_web,
)
progress.end_processing()
@@ -1328,12 +1401,14 @@ def main():
report.reddit = deduped_reddit
report.x = deduped_x
report.youtube = deduped_youtube
report.tiktok = deduped_tiktok
report.hackernews = deduped_hn
report.polymarket = deduped_pm
report.web = deduped_web
report.reddit_error = reddit_error
report.x_error = x_error
report.youtube_error = youtube_error
report.tiktok_error = tiktok_error
report.hackernews_error = hackernews_error
report.polymarket_error = polymarket_error
report.web_error = web_error
@@ -1349,7 +1424,7 @@ def main():
if sources == "web":
progress.show_web_only_complete()
else:
progress.show_complete(len(deduped_reddit), len(deduped_x), len(deduped_youtube), len(deduped_hn), len(deduped_pm))
progress.show_complete(len(deduped_reddit), len(deduped_x), len(deduped_youtube), len(deduped_hn), len(deduped_pm), len(deduped_tiktok))
# Build source info for status footer
source_info = {}
@@ -1364,6 +1439,8 @@ def main():
source_info["youtube_skip_reason"] = "yt-dlp not installed — fix: brew install yt-dlp"
elif has_ytdlp and not report.youtube:
source_info["youtube_skip_reason"] = "0 results (query may be too specific)"
if not has_apify:
source_info["tiktok_skip_reason"] = "No APIFY_API_TOKEN — sign up free at apify.com"
if not web_source:
source_info["web_skip_reason"] = "assistant will use WebSearch (add BRAVE_API_KEY for native search)"
+82
View File
@@ -0,0 +1,82 @@
"""Shared Apify client utilities for last30days sources.
Provides a common wrapper around the apify-client SDK so that
TikTok, Facebook, Instagram (future) all share the same client
initialization, error handling, and cost-control patterns.
One APIFY_API_TOKEN covers all Apify-backed sources.
"""
import sys
from typing import Any, Dict, List, Optional
try:
from apify_client import ApifyClient
except ImportError:
ApifyClient = None
def is_apify_available() -> bool:
"""Check if the apify-client library is installed."""
return ApifyClient is not None
def get_apify_client(token: str) -> "ApifyClient":
"""Initialize Apify client with token.
Args:
token: Apify API token (from https://console.apify.com)
Returns:
Initialized ApifyClient instance
Raises:
ImportError: If apify-client is not installed
"""
if ApifyClient is None:
raise ImportError(
"apify-client is not installed. Run: pip install apify-client"
)
return ApifyClient(token=token)
def run_actor_sync(
client: "ApifyClient",
actor_id: str,
run_input: Dict[str, Any],
timeout_secs: int = 300,
max_items: Optional[int] = None,
) -> List[Dict[str, Any]]:
"""Run an Apify actor synchronously and return dataset items.
Args:
client: Initialized ApifyClient
actor_id: Actor identifier, e.g. "clockworks/tiktok-scraper"
run_input: Actor-specific input dict
timeout_secs: Max wait time (default 5 min)
max_items: Cap on returned items (cost control)
Returns:
List of result dicts from the actor's default dataset
"""
_log(f"Running actor {actor_id} (timeout={timeout_secs}s)")
run = client.actor(actor_id).call(
run_input=run_input,
timeout_secs=timeout_secs,
)
dataset_id = run["defaultDatasetId"]
items = list(client.dataset(dataset_id).iterate_items())
if max_items and len(items) > max_items:
items = items[:max_items]
_log(f"Actor {actor_id} returned {len(items)} items")
return items
def _log(msg: str):
"""Log to stderr."""
sys.stderr.write(f"[Apify] {msg}\n")
sys.stderr.flush()
+13 -1
View File
@@ -45,7 +45,7 @@ def jaccard_similarity(set1: Set[str], set2: Set[str]) -> float:
return intersection / union if union > 0 else 0.0
AnyItem = Union[schema.RedditItem, schema.XItem, schema.YouTubeItem,
AnyItem = Union[schema.RedditItem, schema.XItem, schema.YouTubeItem, schema.TikTokItem,
schema.HackerNewsItem, schema.PolymarketItem, schema.WebSearchItem]
@@ -57,6 +57,8 @@ def get_item_text(item: AnyItem) -> str:
return item.title
elif isinstance(item, schema.YouTubeItem):
return f"{item.title} {item.channel_name}"
elif isinstance(item, schema.TikTokItem):
return f"{item.text} {item.author_name}"
elif isinstance(item, schema.PolymarketItem):
return f"{item.title} {item.question}"
elif isinstance(item, schema.WebSearchItem):
@@ -74,6 +76,8 @@ def _get_cross_source_text(item: AnyItem) -> str:
"""
if isinstance(item, schema.XItem):
return item.text[:100]
if isinstance(item, schema.TikTokItem):
return item.text[:100]
if isinstance(item, schema.HackerNewsItem):
title = item.title
if title.startswith("Show HN:"):
@@ -194,6 +198,14 @@ def dedupe_youtube(
return dedupe_items(items, threshold)
def dedupe_tiktok(
items: List[schema.TikTokItem],
threshold: float = 0.7,
) -> List[schema.TikTokItem]:
"""Dedupe TikTok items."""
return dedupe_items(items, threshold)
def dedupe_hackernews(
items: List[schema.HackerNewsItem],
threshold: float = 0.7,
+10
View File
@@ -203,6 +203,7 @@ def get_config() -> Dict[str, Any]:
('OPENAI_MODEL_PIN', None),
('XAI_MODEL_POLICY', 'latest'),
('XAI_MODEL_PIN', None),
('APIFY_API_TOKEN', None),
('AUTH_TOKEN', None),
('CT0', None),
]
@@ -406,6 +407,15 @@ def is_polymarket_available() -> bool:
return True
def is_apify_available(config: Dict[str, Any]) -> bool:
"""Check if Apify token is configured for TikTok/social scraping.
Returns True if APIFY_API_TOKEN is set. One token covers
TikTok, Facebook, Instagram (all Apify-backed sources).
"""
return bool(config.get('APIFY_API_TOKEN'))
def get_x_source_status(config: Dict[str, Any]) -> Dict[str, Any]:
"""Get detailed X source status for UI decisions.
+48 -1
View File
@@ -4,7 +4,7 @@ from typing import Any, Dict, List, TypeVar, Union
from . import dates, schema
T = TypeVar("T", schema.RedditItem, schema.XItem, schema.WebSearchItem, schema.YouTubeItem, schema.HackerNewsItem, schema.PolymarketItem)
T = TypeVar("T", schema.RedditItem, schema.XItem, schema.WebSearchItem, schema.YouTubeItem, schema.TikTokItem, schema.HackerNewsItem, schema.PolymarketItem)
def filter_by_date_range(
@@ -200,6 +200,53 @@ def normalize_youtube_items(
return normalized
def normalize_tiktok_items(
items: List[Dict[str, Any]],
from_date: str,
to_date: str,
) -> List[schema.TikTokItem]:
"""Normalize raw TikTok items to schema.
Args:
items: Raw TikTok items from Apify
from_date: Start of date range
to_date: End of date range
Returns:
List of TikTokItem objects
"""
normalized = []
for i, item in enumerate(items):
# Parse engagement
eng_raw = item.get("engagement") or {}
engagement = schema.Engagement(
views=eng_raw.get("views"),
likes=eng_raw.get("likes"),
num_comments=eng_raw.get("comments"),
shares=eng_raw.get("shares"),
)
# TikTok dates are reliable (exact timestamps from Apify)
date_str = item.get("date")
normalized.append(schema.TikTokItem(
id=f"TK{i+1}",
text=item.get("text", ""),
url=item.get("url", ""),
author_name=item.get("author_name", ""),
date=date_str,
date_confidence="high",
engagement=engagement,
caption_snippet=item.get("caption_snippet", ""),
hashtags=item.get("hashtags", []),
relevance=item.get("relevance", 0.7),
why_relevant=item.get("why_relevant", ""),
))
return normalized
def normalize_hackernews_items(
items: List[Dict[str, Any]],
from_date: str,
+75 -2
View File
@@ -24,6 +24,8 @@ def _xref_tag(item) -> str:
source_names.add('X')
elif ref_id.startswith('YT'):
source_names.add('YouTube')
elif ref_id.startswith('TK'):
source_names.add('TikTok')
elif ref_id.startswith('HN'):
source_names.add('HN')
elif ref_id.startswith('PM'):
@@ -57,8 +59,10 @@ def _assess_data_freshness(report: schema.Report) -> dict:
hn_recent = sum(1 for h in report.hackernews if h.date and h.date >= report.range_from)
pm_recent = sum(1 for p in report.polymarket if p.date and p.date >= report.range_from)
total_recent = reddit_recent + x_recent + web_recent + hn_recent + pm_recent
total_items = len(report.reddit) + len(report.x) + len(report.web) + len(report.hackernews) + len(report.polymarket)
tiktok_recent = sum(1 for t in report.tiktok if t.date and t.date >= report.range_from)
total_recent = reddit_recent + x_recent + web_recent + hn_recent + pm_recent + tiktok_recent
total_items = len(report.reddit) + len(report.x) + len(report.web) + len(report.hackernews) + len(report.polymarket) + len(report.tiktok)
return {
"reddit_recent": reddit_recent,
@@ -244,6 +248,42 @@ def render_compact(report: schema.Report, limit: int = 15, missing_keys: str = "
lines.append(f" *{item.why_relevant}*")
lines.append("")
# TikTok items
if report.tiktok_error:
lines.append("### TikTok Videos")
lines.append("")
lines.append(f"**ERROR:** {report.tiktok_error}")
lines.append("")
elif report.tiktok:
lines.append("### TikTok Videos")
lines.append("")
for item in report.tiktok[:limit]:
eng_str = ""
if item.engagement:
eng = item.engagement
parts = []
if eng.views is not None:
parts.append(f"{eng.views:,} views")
if eng.likes is not None:
parts.append(f"{eng.likes:,} likes")
if parts:
eng_str = f" [{', '.join(parts)}]"
date_str = f" ({item.date})" if item.date else ""
lines.append(f"**{item.id}** (score:{item.score}) @{item.author_name}{date_str}{eng_str}{_xref_tag(item)}")
lines.append(f" {item.text[:200]}")
lines.append(f" {item.url}")
if item.caption_snippet and item.caption_snippet != item.text[:len(item.caption_snippet)]:
snippet = item.caption_snippet[:200]
if len(item.caption_snippet) > 200:
snippet += "..."
lines.append(f" Caption: {snippet}")
if item.hashtags:
lines.append(f" Tags: {' '.join('#' + h for h in item.hashtags[:8])}")
lines.append(f" *{item.why_relevant}*")
lines.append("")
# Hacker News items
if report.hackernews_error:
lines.append("### Hacker News Stories")
@@ -407,6 +447,14 @@ def render_source_status(report: schema.Report, source_info: dict = None) -> str
lines.append(f" ✅ YouTube: {len(report.youtube)} videos ({with_transcripts} with transcripts)")
# Hide when zero results (no skip reason line needed)
# TikTok
if report.tiktok_error:
lines.append(f" ❌ TikTok: error — {report.tiktok_error}")
elif report.tiktok:
with_captions = sum(1 for v in report.tiktok if getattr(v, 'caption_snippet', None))
lines.append(f" ✅ TikTok: {len(report.tiktok)} videos ({with_captions} with captions)")
# Hide when zero results
# Hacker News
if report.hackernews_error:
lines.append(f" ❌ HN: error - {report.hackernews_error}")
@@ -458,6 +506,8 @@ def render_context_snippet(report: schema.Report) -> str:
all_items.append((item.score, "Reddit", item.title, item.url))
for item in report.x[:5]:
all_items.append((item.score, "X", item.text[:50] + "...", item.url))
for item in report.tiktok[:5]:
all_items.append((item.score, "TikTok", item.text[:50] + "...", item.url))
for item in report.hackernews[:5]:
all_items.append((item.score, "HN", item.title[:50] + "...", item.hn_url))
for item in report.polymarket[:5]:
@@ -551,6 +601,29 @@ def render_full_report(report: schema.Report) -> str:
lines.append(f"> {item.text}")
lines.append("")
# TikTok section
if report.tiktok:
lines.append("## TikTok Videos")
lines.append("")
for item in report.tiktok:
lines.append(f"### {item.id}: @{item.author_name}")
lines.append("")
lines.append(f"- **URL:** {item.url}")
lines.append(f"- **Date:** {item.date or 'Unknown'}")
lines.append(f"- **Score:** {item.score}/100")
lines.append(f"- **Relevance:** {item.why_relevant}")
if item.engagement:
eng = item.engagement
lines.append(f"- **Engagement:** {eng.views or '?'} views, {eng.likes or '?'} likes, {eng.num_comments or '?'} comments")
if item.hashtags:
lines.append(f"- **Hashtags:** {' '.join('#' + h for h in item.hashtags[:10])}")
lines.append("")
lines.append(f"> {item.text[:300]}")
lines.append("")
# HN section
if report.hackernews:
lines.append("## Hacker News Stories")
+75
View File
@@ -22,6 +22,9 @@ class Engagement:
# YouTube fields
views: Optional[int] = None
# TikTok / Facebook fields
shares: Optional[int] = None
# Polymarket fields
volume: Optional[float] = None
liquidity: Optional[float] = None
@@ -44,6 +47,8 @@ class Engagement:
d['quotes'] = self.quotes
if self.views is not None:
d['views'] = self.views
if self.shares is not None:
d['shares'] = self.shares
if self.volume is not None:
d['volume'] = self.volume
if self.liquidity is not None:
@@ -231,6 +236,45 @@ class YouTubeItem:
return d
@dataclass
class TikTokItem:
"""Normalized TikTok item."""
id: str # video_id
text: str # caption/description
url: str # webVideoUrl
author_name: str # authorMeta.name
date: Optional[str] = None
date_confidence: str = "high" # Apify provides exact timestamps
engagement: Optional[Engagement] = None # views, likes, num_comments, shares
caption_snippet: str = "" # spoken-word caption (if available), else text
hashtags: List[str] = field(default_factory=list)
relevance: float = 0.7
why_relevant: str = ""
subs: SubScores = field(default_factory=SubScores)
score: int = 0
cross_refs: List[str] = field(default_factory=list)
def to_dict(self) -> Dict[str, Any]:
d = {
'id': self.id,
'text': self.text,
'url': self.url,
'author_name': self.author_name,
'date': self.date,
'date_confidence': self.date_confidence,
'engagement': self.engagement.to_dict() if self.engagement else None,
'caption_snippet': self.caption_snippet,
'hashtags': self.hashtags,
'relevance': self.relevance,
'why_relevant': self.why_relevant,
'subs': self.subs.to_dict(),
'score': self.score,
}
if self.cross_refs:
d['cross_refs'] = self.cross_refs
return d
@dataclass
class HackerNewsItem:
"""Normalized Hacker News item."""
@@ -329,6 +373,7 @@ class Report:
x: List[XItem] = field(default_factory=list)
web: List[WebSearchItem] = field(default_factory=list)
youtube: List[YouTubeItem] = field(default_factory=list)
tiktok: List[TikTokItem] = field(default_factory=list)
hackernews: List[HackerNewsItem] = field(default_factory=list)
polymarket: List[PolymarketItem] = field(default_factory=list)
best_practices: List[str] = field(default_factory=list)
@@ -339,6 +384,7 @@ class Report:
x_error: Optional[str] = None
web_error: Optional[str] = None
youtube_error: Optional[str] = None
tiktok_error: Optional[str] = None
hackernews_error: Optional[str] = None
polymarket_error: Optional[str] = None
# Handle resolution
@@ -362,6 +408,7 @@ class Report:
'x': [x.to_dict() for x in self.x],
'web': [w.to_dict() for w in self.web],
'youtube': [y.to_dict() for y in self.youtube],
'tiktok': [t.to_dict() for t in self.tiktok],
'hackernews': [h.to_dict() for h in self.hackernews],
'polymarket': [p.to_dict() for p in self.polymarket],
'best_practices': self.best_practices,
@@ -378,6 +425,8 @@ class Report:
d['web_error'] = self.web_error
if self.youtube_error:
d['youtube_error'] = self.youtube_error
if self.tiktok_error:
d['tiktok_error'] = self.tiktok_error
if self.hackernews_error:
d['hackernews_error'] = self.hackernews_error
if self.polymarket_error:
@@ -485,6 +534,30 @@ class Report:
cross_refs=y.get('cross_refs', []),
))
# Reconstruct TikTok items
tiktok_items = []
for t in data.get('tiktok', []):
eng = None
if t.get('engagement'):
eng = Engagement(**t['engagement'])
subs = SubScores(**t.get('subs', {})) if t.get('subs') else SubScores()
tiktok_items.append(TikTokItem(
id=t['id'],
text=t.get('text', ''),
url=t['url'],
author_name=t.get('author_name', ''),
date=t.get('date'),
date_confidence=t.get('date_confidence', 'high'),
engagement=eng,
caption_snippet=t.get('caption_snippet', ''),
hashtags=t.get('hashtags', []),
relevance=t.get('relevance', 0.7),
why_relevant=t.get('why_relevant', ''),
subs=subs,
score=t.get('score', 0),
cross_refs=t.get('cross_refs', []),
))
# Reconstruct HackerNews items
hn_items = []
for h in data.get('hackernews', []):
@@ -549,6 +622,7 @@ class Report:
x=x_items,
web=web_items,
youtube=youtube_items,
tiktok=tiktok_items,
hackernews=hn_items,
polymarket=pm_items,
best_practices=data.get('best_practices', []),
@@ -558,6 +632,7 @@ class Report:
x_error=data.get('x_error'),
web_error=data.get('web_error'),
youtube_error=data.get('youtube_error'),
tiktok_error=data.get('tiktok_error'),
hackernews_error=data.get('hackernews_error'),
polymarket_error=data.get('polymarket_error'),
resolved_x_handle=data.get('resolved_x_handle'),
+66 -5
View File
@@ -280,6 +280,65 @@ def score_youtube_items(items: List[schema.YouTubeItem]) -> List[schema.YouTubeI
return items
def compute_tiktok_engagement_raw(engagement: Optional[schema.Engagement]) -> Optional[float]:
"""Compute raw engagement score for TikTok item.
Formula: 0.50*log1p(views) + 0.30*log1p(likes) + 0.20*log1p(comments)
Views dominate on TikTok they're the primary discovery signal.
"""
if engagement is None:
return None
if engagement.views is None and engagement.likes is None:
return None
views = log1p_safe(engagement.views)
likes = log1p_safe(engagement.likes)
comments = log1p_safe(engagement.num_comments)
return 0.50 * views + 0.30 * likes + 0.20 * comments
def score_tiktok_items(items: List[schema.TikTokItem]) -> List[schema.TikTokItem]:
"""Compute scores for TikTok items.
Uses same weight structure as YouTube (relevance + recency + engagement).
"""
if not items:
return items
eng_raw = [compute_tiktok_engagement_raw(item.engagement) for item in items]
eng_normalized = normalize_to_100(eng_raw)
for i, item in enumerate(items):
rel_score = int(item.relevance * 100)
rec_score = dates.recency_score(item.date)
if eng_normalized[i] is not None:
eng_score = int(eng_normalized[i])
else:
eng_score = DEFAULT_ENGAGEMENT
item.subs = schema.SubScores(
relevance=rel_score,
recency=rec_score,
engagement=eng_score,
)
overall = (
WEIGHT_RELEVANCE * rel_score +
WEIGHT_RECENCY * rec_score +
WEIGHT_ENGAGEMENT * eng_score
)
if eng_raw[i] is None:
overall -= UNKNOWN_ENGAGEMENT_PENALTY
item.score = max(0, min(100, int(overall)))
return items
def compute_hackernews_engagement_raw(engagement: Optional[schema.Engagement]) -> Optional[float]:
"""Compute raw engagement score for Hacker News item.
@@ -453,7 +512,7 @@ def score_websearch_items(items: List[schema.WebSearchItem]) -> List[schema.WebS
return items
def sort_items(items: List[Union[schema.RedditItem, schema.XItem, schema.WebSearchItem, schema.YouTubeItem, schema.HackerNewsItem, schema.PolymarketItem]]) -> List:
def sort_items(items: List[Union[schema.RedditItem, schema.XItem, schema.WebSearchItem, schema.YouTubeItem, schema.TikTokItem, schema.HackerNewsItem, schema.PolymarketItem]]) -> List:
"""Sort items by score (descending), then date, then source priority.
Args:
@@ -470,19 +529,21 @@ def sort_items(items: List[Union[schema.RedditItem, schema.XItem, schema.WebSear
date = item.date or "0000-00-00"
date_key = -int(date.replace("-", ""))
# Tertiary: source priority (Reddit > X > YouTube > HN > Polymarket > WebSearch)
# Tertiary: source priority (Reddit > X > YouTube > TikTok > HN > Polymarket > WebSearch)
if isinstance(item, schema.RedditItem):
source_priority = 0
elif isinstance(item, schema.XItem):
source_priority = 1
elif isinstance(item, schema.YouTubeItem):
source_priority = 2
elif isinstance(item, schema.HackerNewsItem):
elif isinstance(item, schema.TikTokItem):
source_priority = 3
elif isinstance(item, schema.PolymarketItem):
elif isinstance(item, schema.HackerNewsItem):
source_priority = 4
else: # WebSearchItem
elif isinstance(item, schema.PolymarketItem):
source_priority = 5
else: # WebSearchItem
source_priority = 6
# Quaternary: title/text for stability
text = getattr(item, "title", "") or getattr(item, "text", "")
+385
View File
@@ -0,0 +1,385 @@
"""TikTok search via Apify clockworks/tiktok-scraper for /last30days.
Uses the Apify platform to search TikTok by keyword, extract engagement
metrics (views, likes, comments), and optionally pull video captions.
Requires APIFY_API_TOKEN in config. Free tier: $5/month credits.
"""
import re
import sys
from datetime import datetime, timezone
from typing import Any, Dict, List, Optional, Set
from . import apify_client_wrapper
ACTOR_ID = "clockworks/tiktok-scraper"
# Depth configurations: how many results to fetch / captions to extract
DEPTH_CONFIG = {
"quick": {"results_per_page": 10, "max_captions": 3},
"default": {"results_per_page": 20, "max_captions": 5},
"deep": {"results_per_page": 40, "max_captions": 8},
}
# Max words to keep from each caption
CAPTION_MAX_WORDS = 500
# Stopwords for relevance computation (shared with youtube_yt.py pattern)
STOPWORDS = frozenset({
'the', 'a', 'an', 'to', 'for', 'how', 'is', 'in', 'of', 'on',
'and', 'with', 'from', 'by', 'at', 'this', 'that', 'it', 'my',
'your', 'i', 'me', 'we', 'you', 'what', 'are', 'do', 'can',
'its', 'be', 'or', 'not', 'no', 'so', 'if', 'but', 'about',
'all', 'just', 'get', 'has', 'have', 'was', 'will',
})
# Synonym groups for relevance scoring
SYNONYMS = {
'hip': {'rap', 'hiphop'},
'hop': {'rap', 'hiphop'},
'rap': {'hip', 'hop', 'hiphop'},
'hiphop': {'rap', 'hip', 'hop'},
'js': {'javascript'},
'javascript': {'js'},
'ts': {'typescript'},
'typescript': {'ts'},
'ai': {'artificial', 'intelligence'},
'ml': {'machine', 'learning'},
'react': {'reactjs'},
'reactjs': {'react'},
}
def _tokenize(text: str) -> Set[str]:
"""Lowercase, strip punctuation, remove stopwords, drop single-char tokens."""
words = re.sub(r'[^\w\s]', ' ', text.lower()).split()
tokens = {w for w in words if w not in STOPWORDS and len(w) > 1}
expanded = set(tokens)
for t in tokens:
if t in SYNONYMS:
expanded.update(SYNONYMS[t])
return expanded
def _compute_relevance(query: str, text: str, hashtags: List[str] = None) -> float:
"""Compute relevance as ratio of query tokens found in text + hashtags.
Uses ratio overlap (intersection / query_length). Hashtags provide
a TikTok-specific relevance boost. Floors at 0.1.
"""
q_tokens = _tokenize(query)
# Combine text and hashtags for matching
combined = text
if hashtags:
combined = f"{text} {' '.join(hashtags)}"
t_tokens = _tokenize(combined)
# Split concatenated hashtags (e.g., "claudecode" → "claude", "code")
if hashtags:
for tag in hashtags:
tag_lower = tag.lower()
for qt in q_tokens:
if qt in tag_lower and qt != tag_lower:
t_tokens.add(qt)
if not q_tokens:
return 0.5 # Neutral fallback
overlap = len(q_tokens & t_tokens)
ratio = overlap / len(q_tokens)
return max(0.1, min(1.0, ratio))
def _extract_core_subject(topic: str) -> str:
"""Extract core subject from verbose query for TikTok search.
Strips meta/research words to keep only the core product/concept name.
"""
text = topic.lower().strip()
# Strip multi-word prefixes
prefixes = [
'what are the best', 'what is the best', 'what are the latest',
'what are people saying about', 'what do people think about',
'how do i use', 'how to use', 'how to',
'what are', 'what is', 'tips for', 'best practices for',
]
for p in prefixes:
if text.startswith(p + ' '):
text = text[len(p):].strip()
# Strip individual noise words
noise = {
'best', 'top', 'good', 'great', 'awesome', 'killer',
'latest', 'new', 'news', 'update', 'updates',
'trending', 'hottest', 'popular', 'viral',
'practices', 'features',
'recommendations', 'advice',
'prompt', 'prompts', 'prompting',
'methods', 'strategies', 'approaches',
}
words = text.split()
filtered = [w for w in words if w not in noise]
result = ' '.join(filtered) if filtered else text
return result.rstrip('?!.')
def _log(msg: str):
"""Log to stderr."""
sys.stderr.write(f"[TikTok] {msg}\n")
sys.stderr.flush()
def _parse_date(item: Dict[str, Any]) -> Optional[str]:
"""Parse date from Apify TikTok item to YYYY-MM-DD.
Handles both createTimeISO (ISO string) and createTime (unix timestamp).
"""
iso = item.get("createTimeISO")
if iso:
try:
dt = datetime.fromisoformat(iso.replace("Z", "+00:00"))
return dt.strftime("%Y-%m-%d")
except (ValueError, TypeError):
pass
ts = item.get("createTime")
if ts:
try:
dt = datetime.fromtimestamp(int(ts), tz=timezone.utc)
return dt.strftime("%Y-%m-%d")
except (ValueError, TypeError, OSError):
pass
return None
def search_tiktok(
topic: str,
from_date: str,
to_date: str,
depth: str = "default",
token: str = None,
) -> Dict[str, Any]:
"""Search TikTok via Apify.
Args:
topic: Search topic
from_date: Start date (YYYY-MM-DD)
to_date: End date (YYYY-MM-DD)
depth: 'quick', 'default', or 'deep'
token: Apify API token
Returns:
Dict with 'items' list and optional 'error'.
"""
if not token:
return {"items": [], "error": "No APIFY_API_TOKEN configured"}
if not apify_client_wrapper.is_apify_available():
return {"items": [], "error": "apify-client not installed (pip install apify-client)"}
config = DEPTH_CONFIG.get(depth, DEPTH_CONFIG["default"])
core_topic = _extract_core_subject(topic)
_log(f"Searching TikTok for '{core_topic}' (depth={depth}, count={config['results_per_page']})")
try:
client = apify_client_wrapper.get_apify_client(token)
run_input = {
"searchQueries": [core_topic],
"resultsPerPage": config["results_per_page"],
"shouldDownloadSubtitles": False,
"shouldDownloadVideos": False,
"shouldDownloadCovers": False,
}
raw_items = apify_client_wrapper.run_actor_sync(
client, ACTOR_ID, run_input,
timeout_secs=120,
max_items=config["results_per_page"],
)
except Exception as e:
_log(f"Apify error: {e}")
return {"items": [], "error": f"{type(e).__name__}: {e}"}
# Parse items
items = []
for raw in raw_items:
video_id = str(raw.get("id", ""))
text = raw.get("text", "")
play_count = raw.get("playCount") or 0
digg_count = raw.get("diggCount") or 0
comment_count = raw.get("commentCount") or 0
share_count = raw.get("shareCount") or 0
author_meta = raw.get("authorMeta") or {}
author_name = author_meta.get("name", "")
web_url = raw.get("webVideoUrl", "")
hashtags_raw = raw.get("hashtags") or []
hashtag_names = [h.get("name", "") for h in hashtags_raw if isinstance(h, dict)]
duration = (raw.get("videoMeta") or {}).get("duration")
date_str = _parse_date(raw)
# Compute relevance with hashtag boost
relevance = _compute_relevance(core_topic, text, hashtag_names)
items.append({
"video_id": video_id,
"text": text,
"url": web_url or f"https://www.tiktok.com/@{author_name}/video/{video_id}",
"author_name": author_name,
"date": date_str,
"engagement": {
"views": play_count,
"likes": digg_count,
"comments": comment_count,
"shares": share_count,
},
"hashtags": hashtag_names,
"duration": duration,
"relevance": relevance,
"why_relevant": f"TikTok: {text[:60]}" if text else f"TikTok: {core_topic}",
"caption_snippet": "", # populated by fetch_captions
})
# Hard date filter
in_range = [i for i in items if i["date"] and from_date <= i["date"] <= to_date]
out_of_range = len(items) - len(in_range)
if in_range:
items = in_range
if out_of_range:
_log(f"Filtered {out_of_range} videos outside date range")
else:
_log(f"No videos within date range, keeping all {len(items)}")
# Sort by views descending
items.sort(key=lambda x: x["engagement"]["views"], reverse=True)
_log(f"Found {len(items)} TikTok videos")
return {"items": items}
def fetch_captions(
video_items: List[Dict[str, Any]],
token: str,
depth: str = "default",
) -> Dict[str, str]:
"""Fetch captions for top N TikTok videos.
Strategy:
1. Primary: Use the 'text' field (video description) always free
2. For top N, re-run actor with shouldDownloadSubtitles for spoken-word
Args:
video_items: Items from search_tiktok()
token: Apify API token
depth: Depth level for caption limit
Returns:
Dict mapping video_id caption text (truncated to 500 words)
"""
config = DEPTH_CONFIG.get(depth, DEPTH_CONFIG["default"])
max_captions = config["max_captions"]
if not video_items or not token:
return {}
top_items = video_items[:max_captions]
_log(f"Enriching captions for {len(top_items)} videos")
captions = {}
# First pass: use text field as caption (always available, free)
for item in top_items:
vid = item["video_id"]
text = item.get("text", "")
if text:
words = text.split()
if len(words) > CAPTION_MAX_WORDS:
text = ' '.join(words[:CAPTION_MAX_WORDS]) + '...'
captions[vid] = text
# Second pass: try to get spoken-word subtitles for top videos
try:
urls = [item["url"] for item in top_items if item.get("url")]
if urls:
client = apify_client_wrapper.get_apify_client(token)
run_input = {
"postURLs": urls,
"shouldDownloadSubtitles": True,
"shouldDownloadVideos": False,
"shouldDownloadCovers": False,
}
subtitle_items = apify_client_wrapper.run_actor_sync(
client, ACTOR_ID, run_input,
timeout_secs=60,
max_items=max_captions,
)
for raw in subtitle_items:
vid = str(raw.get("id", ""))
# Check for subtitle text in the response
subtitle_text = raw.get("subtitleText") or raw.get("subtitles") or ""
if isinstance(subtitle_text, list):
subtitle_text = " ".join(str(s) for s in subtitle_text)
if subtitle_text and vid:
words = subtitle_text.split()
if len(words) > CAPTION_MAX_WORDS:
subtitle_text = ' '.join(words[:CAPTION_MAX_WORDS]) + '...'
captions[vid] = subtitle_text # Override text with spoken-word
except Exception as e:
_log(f"Subtitle enrichment failed (using text captions): {e}")
got = sum(1 for v in captions.values() if v)
_log(f"Got captions for {got}/{len(top_items)} videos")
return captions
def search_and_enrich(
topic: str,
from_date: str,
to_date: str,
depth: str = "default",
token: str = None,
) -> Dict[str, Any]:
"""Full TikTok search: find videos, then fetch captions for top results.
Args:
topic: Search topic
from_date: Start date (YYYY-MM-DD)
to_date: End date (YYYY-MM-DD)
depth: 'quick', 'default', or 'deep'
token: Apify API token
Returns:
Dict with 'items' list. Each item has a 'caption_snippet' field.
"""
# Step 1: Search
search_result = search_tiktok(topic, from_date, to_date, depth, token)
items = search_result.get("items", [])
if not items:
return search_result
# Step 2: Fetch captions for top N
captions = fetch_captions(items, token, depth)
# Step 3: Attach captions to items
for item in items:
vid = item["video_id"]
caption = captions.get(vid)
if caption:
item["caption_snippet"] = caption
return {"items": items, "error": search_result.get("error")}
def parse_tiktok_response(response: Dict[str, Any]) -> List[Dict[str, Any]]:
"""Parse TikTok search response to normalized format.
Returns:
List of item dicts ready for normalization.
"""
return response.get("items", [])
+20 -1
View File
@@ -71,6 +71,12 @@ YOUTUBE_MESSAGES = [
"Fetching transcripts...",
]
TIKTOK_MESSAGES = [
"Searching TikTok for trending videos...",
"Finding what's viral on TikTok...",
"Scanning TikTok for relevant content...",
]
HN_MESSAGES = [
"Searching Hacker News...",
"Scanning HN front page stories...",
@@ -271,6 +277,15 @@ class ProgressDisplay:
if self.spinner:
self.spinner.stop(f"{Colors.RED}YouTube{Colors.RESET} Found {count} videos")
def start_tiktok(self):
msg = random.choice(TIKTOK_MESSAGES)
self.spinner = Spinner(f"{Colors.PURPLE}TikTok{Colors.RESET} {msg}", Colors.PURPLE, quiet=True)
self.spinner.start()
def end_tiktok(self, count: int):
if self.spinner:
self.spinner.stop(f"{Colors.PURPLE}TikTok{Colors.RESET} Found {count} videos")
def start_hackernews(self):
msg = random.choice(HN_MESSAGES)
self.spinner = Spinner(f"{Colors.YELLOW}HN{Colors.RESET} {msg}", Colors.YELLOW, quiet=True)
@@ -298,7 +313,7 @@ class ProgressDisplay:
if self.spinner:
self.spinner.stop()
def show_complete(self, reddit_count: int, x_count: int, youtube_count: int = 0, hn_count: int = 0, pm_count: int = 0):
def show_complete(self, reddit_count: int, x_count: int, youtube_count: int = 0, hn_count: int = 0, pm_count: int = 0, tiktok_count: int = 0):
elapsed = time.time() - self.start_time
if IS_TTY:
sys.stderr.write(f"\n{Colors.GREEN}{Colors.BOLD}✓ Research complete{Colors.RESET} ")
@@ -307,6 +322,8 @@ class ProgressDisplay:
sys.stderr.write(f"{Colors.CYAN}X:{Colors.RESET} {x_count} posts")
if youtube_count:
sys.stderr.write(f" {Colors.RED}YouTube:{Colors.RESET} {youtube_count} videos")
if tiktok_count:
sys.stderr.write(f" {Colors.PURPLE}TikTok:{Colors.RESET} {tiktok_count} videos")
if hn_count:
sys.stderr.write(f" {Colors.YELLOW}HN:{Colors.RESET} {hn_count} stories")
if pm_count:
@@ -316,6 +333,8 @@ class ProgressDisplay:
parts = [f"Reddit: {reddit_count} threads", f"X: {x_count} posts"]
if youtube_count:
parts.append(f"YouTube: {youtube_count} videos")
if tiktok_count:
parts.append(f"TikTok: {tiktok_count} videos")
if hn_count:
parts.append(f"HN: {hn_count} stories")
if pm_count:
+259
View File
@@ -0,0 +1,259 @@
"""Tests for TikTok module (search, normalize, score, dedupe, render)."""
import json
import sys
import unittest
from pathlib import Path
# Add lib to path
sys.path.insert(0, str(Path(__file__).parent.parent / "scripts"))
from lib import schema, score, normalize, dedupe, render
from lib import tiktok
class TestTikTokRelevance(unittest.TestCase):
"""Test relevance scoring for TikTok items."""
def test_exact_match(self):
rel = tiktok._compute_relevance("claude code", "Claude Code tricks and tips")
self.assertGreaterEqual(rel, 0.8)
def test_partial_match(self):
rel = tiktok._compute_relevance("claude code tips", "Best AI tools for coding")
self.assertLess(rel, 0.5)
def test_hashtag_boost(self):
"""Hashtags should boost relevance."""
rel_no_hash = tiktok._compute_relevance("claude code", "random video about stuff")
rel_with_hash = tiktok._compute_relevance("claude code", "random video about stuff", ["claudecode", "ai"])
self.assertGreater(rel_with_hash, rel_no_hash)
def test_empty_query(self):
rel = tiktok._compute_relevance("", "Some video title")
self.assertEqual(rel, 0.5)
def test_floor(self):
rel = tiktok._compute_relevance("quantum physics", "cat dancing video")
self.assertGreaterEqual(rel, 0.1)
class TestExtractCoreSubject(unittest.TestCase):
"""Test core subject extraction for TikTok search."""
def test_strips_prefix(self):
result = tiktok._extract_core_subject("what are the best claude code tips")
self.assertNotIn("what are the best", result)
self.assertIn("claude", result)
def test_strips_noise(self):
result = tiktok._extract_core_subject("latest trending updates on React")
self.assertNotIn("latest", result)
self.assertNotIn("trending", result)
self.assertIn("react", result.lower())
def test_preserves_core(self):
result = tiktok._extract_core_subject("Claude Code")
self.assertEqual(result, "claude code")
class TestParseDate(unittest.TestCase):
"""Test date parsing from Apify items."""
def test_iso_date(self):
item = {"createTimeISO": "2026-02-28T17:44:35.000Z"}
self.assertEqual(tiktok._parse_date(item), "2026-02-28")
def test_unix_timestamp(self):
item = {"createTime": 1756403075}
result = tiktok._parse_date(item)
self.assertIsNotNone(result)
self.assertRegex(result, r"\d{4}-\d{2}-\d{2}")
def test_no_date(self):
item = {}
self.assertIsNone(tiktok._parse_date(item))
class TestNormalizeTikTokItems(unittest.TestCase):
"""Test TikTok normalization."""
def setUp(self):
self.fixtures_dir = Path(__file__).parent.parent / "fixtures"
with open(self.fixtures_dir / "tiktok_search.json") as f:
data = json.load(f)
self.raw_items = data["items"]
def test_normalizes_items(self):
items = normalize.normalize_tiktok_items(self.raw_items, "2026-02-01", "2026-03-03")
self.assertEqual(len(items), 3)
self.assertIsInstance(items[0], schema.TikTokItem)
def test_ids_are_sequential(self):
items = normalize.normalize_tiktok_items(self.raw_items, "2026-02-01", "2026-03-03")
self.assertEqual(items[0].id, "TK1")
self.assertEqual(items[1].id, "TK2")
self.assertEqual(items[2].id, "TK3")
def test_engagement_parsed(self):
items = normalize.normalize_tiktok_items(self.raw_items, "2026-02-01", "2026-03-03")
eng = items[0].engagement
self.assertIsNotNone(eng)
self.assertEqual(eng.views, 2100000)
self.assertEqual(eng.likes, 45000)
self.assertEqual(eng.shares, 8400)
def test_hashtags_preserved(self):
items = normalize.normalize_tiktok_items(self.raw_items, "2026-02-01", "2026-03-03")
self.assertEqual(items[0].hashtags, ["claudecode", "ai", "coding"])
def test_caption_snippet_preserved(self):
items = normalize.normalize_tiktok_items(self.raw_items, "2026-02-01", "2026-03-03")
self.assertIn("slash commands", items[0].caption_snippet)
class TestScoreTikTokItems(unittest.TestCase):
"""Test TikTok scoring."""
def test_engagement_scoring(self):
eng = schema.Engagement(views=1000000, likes=50000, num_comments=2000)
raw = score.compute_tiktok_engagement_raw(eng)
self.assertIsNotNone(raw)
self.assertGreater(raw, 0)
def test_none_engagement(self):
raw = score.compute_tiktok_engagement_raw(None)
self.assertIsNone(raw)
def test_empty_engagement(self):
eng = schema.Engagement()
raw = score.compute_tiktok_engagement_raw(eng)
self.assertIsNone(raw)
def test_scoring_pipeline(self):
items = [
schema.TikTokItem(
id="TK1", text="High views video", url="https://tiktok.com/1",
author_name="creator1", date="2026-03-01",
engagement=schema.Engagement(views=2000000, likes=50000, num_comments=1000),
relevance=0.9,
),
schema.TikTokItem(
id="TK2", text="Low views video", url="https://tiktok.com/2",
author_name="creator2", date="2026-02-20",
engagement=schema.Engagement(views=1000, likes=50, num_comments=5),
relevance=0.5,
),
]
scored = score.score_tiktok_items(items)
self.assertEqual(len(scored), 2)
self.assertGreater(scored[0].score, 0)
self.assertGreater(scored[0].score, scored[1].score)
class TestDedupeTikTok(unittest.TestCase):
"""Test TikTok deduplication."""
def test_no_dupes(self):
items = [
schema.TikTokItem(id="TK1", text="Totally different video A",
url="https://tiktok.com/1", author_name="a", score=80),
schema.TikTokItem(id="TK2", text="Completely unique video B",
url="https://tiktok.com/2", author_name="b", score=70),
]
result = dedupe.dedupe_tiktok(items)
self.assertEqual(len(result), 2)
def test_removes_dupes(self):
items = [
schema.TikTokItem(id="TK1", text="Claude Code is amazing for AI coding",
url="https://tiktok.com/1", author_name="a", score=80),
schema.TikTokItem(id="TK2", text="Claude Code is amazing for AI coding wow",
url="https://tiktok.com/2", author_name="a", score=60),
]
result = dedupe.dedupe_tiktok(items)
self.assertEqual(len(result), 1)
self.assertEqual(result[0].id, "TK1") # Higher score kept
class TestRenderTikTok(unittest.TestCase):
"""Test TikTok rendering in reports."""
def test_renders_tiktok_section(self):
report = schema.Report(
topic="test", range_from="2026-02-01", range_to="2026-03-03",
generated_at="2026-03-03T00:00:00Z", mode="all",
tiktok=[
schema.TikTokItem(
id="TK1", text="Video caption here", url="https://tiktok.com/1",
author_name="creator", date="2026-03-01", score=85,
engagement=schema.Engagement(views=1000000, likes=50000),
hashtags=["ai", "coding"],
why_relevant="TikTok: Video caption here",
),
],
)
output = render.render_compact(report)
self.assertIn("### TikTok Videos", output)
self.assertIn("TK1", output)
self.assertIn("@creator", output)
self.assertIn("1,000,000 views", output)
def test_renders_source_status(self):
report = schema.Report(
topic="test", range_from="2026-02-01", range_to="2026-03-03",
generated_at="2026-03-03T00:00:00Z", mode="all",
tiktok=[
schema.TikTokItem(
id="TK1", text="test", url="https://tiktok.com/1",
author_name="creator", caption_snippet="some caption",
),
],
)
status = render.render_source_status(report)
self.assertIn("TikTok", status)
self.assertIn("1 videos", status)
def test_xref_tag_tiktok(self):
"""Test that TK prefix is recognized in cross-ref tags."""
item = schema.RedditItem(id="R1", title="test", url="test", subreddit="test",
cross_refs=["TK1"])
tag = render._xref_tag(item)
self.assertIn("TikTok", tag)
class TestSchemaRoundtrip(unittest.TestCase):
"""Test TikTokItem serialization round-trip via Report."""
def test_to_dict_and_back(self):
original = schema.TikTokItem(
id="TK1", text="Test caption", url="https://tiktok.com/1",
author_name="creator", date="2026-03-01",
date_confidence="high",
engagement=schema.Engagement(views=100, likes=10, num_comments=5, shares=3),
caption_snippet="spoken words",
hashtags=["test", "ai"],
relevance=0.8, why_relevant="TikTok: Test",
subs=schema.SubScores(relevance=80, recency=90, engagement=70),
score=80, cross_refs=["R1"],
)
report = schema.Report(
topic="test", range_from="2026-02-01", range_to="2026-03-03",
generated_at="2026-03-03T00:00:00Z", mode="all",
tiktok=[original],
)
d = report.to_dict()
restored = schema.Report.from_dict(d)
self.assertEqual(len(restored.tiktok), 1)
tk = restored.tiktok[0]
self.assertEqual(tk.id, "TK1")
self.assertEqual(tk.author_name, "creator")
self.assertEqual(tk.hashtags, ["test", "ai"])
self.assertEqual(tk.engagement.views, 100)
self.assertEqual(tk.engagement.shares, 3)
self.assertEqual(tk.caption_snippet, "spoken words")
self.assertEqual(tk.cross_refs, ["R1"])
if __name__ == "__main__":
unittest.main()