feat: surface YouTube + TikTok top comments alongside Reddit (#260)
* feat(normalize): pass YouTube top_comments through with Reddit-compatible shape
_normalize_youtube silently dropped top_comments after enrich_with_comments
populated them, so the downstream signals/render/entity layers never saw
YouTube comments. Map likes->score and text->excerpt so the existing
Reddit-compatible readers Just Work.
Shared _remap_comments helper will be reused for TikTok in a later commit.
* feat(tiktok): fetch top comments via ScrapeCreators when opted in
Mirrors the youtube_comments pattern: new env.is_tiktok_comments_available
gate (requires SCRAPECREATORS_API_KEY + tiktok_comments in INCLUDE_SOURCES),
tiktok.enrich_with_comments ranks posts and fetches via
GET /v1/tiktok/video/comments. Vote field is digg_count; text and user.nickname
come across verbatim. Pipeline calls the enricher right after TikTok search
when the gate is open.
Comment-fetch errors never crash the pipeline — the enricher returns an
empty list on 4xx/5xx.
* feat(normalize): pass TikTok top_comments through with digg_count->score mapping
Instagram uses the same shortform normalizer and has no comment fetcher
today, so the key is harmlessly absent there — no Instagram regression.
* feat(signals): add YouTube + TikTok top-comment score to engagement formula
Mirrors Reddit's 10% top-comment slot. Without top_comments present, the
formula reduces to views-dominant weighting; with a high-signal comment,
the item gets a meaningful bump (log1p(10k) ~ 9.2, weighted 0.10 = ~0.92
on the engagement score).
Updated the existing dominant-weight and missing-fields tests to the new
weights (0.45/0.32/0.13 for YT, 0.45/0.27/0.18 for TT). Views still dominate.
* feat(render): source-aware thresholds and vote labels for top comments
10 upvotes on Reddit signals community interest; 10 likes on a viral
TikTok is noise. Introduce per-source minimums (reddit 10, youtube 50,
tiktok 500) and native vote labels ('upvotes' for Reddit, 'likes' for
YT/TT). First-pass numbers — tune after live observation.
* docs: generalize top-comment quoting to YouTube + TikTok, add tiktok_comments opt-in
Synthesis instructions previously called out Reddit top comments only.
Now cover Reddit/YouTube/TikTok uniformly with source-appropriate vote
labels (upvotes vs likes), and explicitly frame YT transcript highlights
and comments as complementary signals. README and setup-wizard copy
document the new tiktok_comments INCLUDE_SOURCES token.
---------
Co-authored-by: Matt Van Horn <455140+mvanhorn@users.noreply.github.com>
This commit is contained in:
+36
-5
@@ -152,13 +152,14 @@ def render_full(report: schema.Report) -> str:
|
||||
lines.append(f" *{item.container}*")
|
||||
if item.snippet:
|
||||
lines.append(f" {item.snippet[:500]}")
|
||||
# Top comments for Reddit
|
||||
# Top comments for Reddit, YouTube, TikTok, HackerNews.
|
||||
top_comments = item.metadata.get("top_comments", [])
|
||||
if top_comments and isinstance(top_comments[0], dict):
|
||||
vote_label = _vote_label_for(item.source)
|
||||
for tc in top_comments[:3]:
|
||||
excerpt = tc.get("excerpt", tc.get("text", ""))[:200]
|
||||
tc_score = tc.get("score", "")
|
||||
lines.append(f" Top comment ({tc_score} upvotes): {excerpt}")
|
||||
lines.append(f" Top comment ({tc_score} {vote_label}): {excerpt}")
|
||||
# Comment insights for Reddit
|
||||
insights = item.metadata.get("comment_insights", [])
|
||||
if insights:
|
||||
@@ -276,7 +277,8 @@ def _render_candidate(candidate: schema.Candidate, prefix: str) -> list[str]:
|
||||
for tc in _top_comments_list(primary):
|
||||
excerpt = tc.get("excerpt") or tc.get("text") or ""
|
||||
score = tc.get("score", "")
|
||||
lines.append(f" - Comment ({score} upvotes): {_truncate(excerpt.strip(), 240)}")
|
||||
vote_label = _vote_label_for(primary.source) if primary else "upvotes"
|
||||
lines.append(f" - Comment ({score} {vote_label}): {_truncate(excerpt.strip(), 240)}")
|
||||
insight = _comment_insight(primary)
|
||||
if insight:
|
||||
lines.append(f" - Insight: {_truncate(insight, 220)}")
|
||||
@@ -582,13 +584,42 @@ def _format_explanation(candidate: schema.Candidate) -> str | None:
|
||||
return candidate.explanation
|
||||
|
||||
|
||||
def _top_comments_list(item: schema.SourceItem | None, limit: int = 3, min_score: int = 10) -> list[dict]:
|
||||
"""Return up to `limit` top comments with score >= min_score."""
|
||||
# Per-source minimum vote counts for showing a top comment in compact emit.
|
||||
# Reddit upvotes, YouTube likes, and TikTok likes are not comparable units —
|
||||
# 10 upvotes on Reddit signals genuine community interest, 10 likes on a
|
||||
# viral TikTok is noise. First-pass values; tune after live observation.
|
||||
_TOP_COMMENT_MIN_SCORE: dict[str, int] = {
|
||||
"reddit": 10,
|
||||
"youtube": 50,
|
||||
"tiktok": 500,
|
||||
"hackernews": 5,
|
||||
}
|
||||
_TOP_COMMENT_VOTE_LABEL: dict[str, str] = {
|
||||
"reddit": "upvotes",
|
||||
"hackernews": "points",
|
||||
"youtube": "likes",
|
||||
"tiktok": "likes",
|
||||
}
|
||||
|
||||
|
||||
def _vote_label_for(source: str) -> str:
|
||||
return _TOP_COMMENT_VOTE_LABEL.get(source, "votes")
|
||||
|
||||
|
||||
def _top_comments_list(item: schema.SourceItem | None, limit: int = 3, min_score: int | None = None) -> list[dict]:
|
||||
"""Return up to `limit` top comments with score at or above the source's minimum.
|
||||
|
||||
If `min_score` is passed explicitly it overrides the per-source default;
|
||||
otherwise the source-keyed map is consulted, with an effective default of 0
|
||||
(always show) for unknown sources so new sources don't get silently hidden.
|
||||
"""
|
||||
if not item:
|
||||
return []
|
||||
comments = item.metadata.get("top_comments") or []
|
||||
if not comments or not isinstance(comments[0], dict):
|
||||
return []
|
||||
if min_score is None:
|
||||
min_score = _TOP_COMMENT_MIN_SCORE.get(item.source, 0)
|
||||
return [c for c in comments if (c.get("score") or 0) >= min_score][:limit]
|
||||
|
||||
|
||||
|
||||
Reference in New Issue
Block a user