Ran 15-way blinded comparison (5 topics x 3 versions). CROSS won all 5 topics (4.74/5.0 avg vs HN 4.10, Base 3.73). Then improved CROSS further: - dedupe.py: hybrid similarity (token+trigram Jaccard) at 0.40 threshold, cross-source links went from 3 to 13 items across 5 topics - render.py: [xref: HN5, HN4] -> [also on: HN, Reddit] for human-readable tags - youtube_yt.py: SYNONYMS dict so "hip hop" matches "rap" (0.33 -> 0.71 score) - SKILL.md: instruction #7 tells Claude to lead with cross-platform signals Validation: improved CROSS scores 4.38/5.0 vs original 3.98 (+0.40), wins 4/5 topics. Biggest gains in specificity (+0.8) and format compliance (+1.0). Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
3.3 KiB
Synthesis Instructions (HN version - with Hacker News)
Judge Agent: Synthesize All Sources
After all searches complete, internally synthesize (don't display stats yet):
- Weight Reddit/X sources HIGHER (they have engagement signals: upvotes, likes)
- Weight YouTube sources HIGH (they have views, likes, and transcript content)
- Weight WebSearch sources LOWER (no engagement data)
- Identify patterns that appear across ALL sources (strongest signals)
- Note any contradictions between sources
- Extract the top 3-5 actionable insights
Internalize the Research
CRITICAL: Ground your synthesis in the ACTUAL research content, not your pre-existing knowledge.
Read the research output carefully. Pay attention to:
- Exact product/tool names mentioned
- Specific quotes and insights from the sources - use THESE, not generic knowledge
- What the sources actually say, not what you assume the topic is about
Citation Rules
CITATION RULE: Cite sources sparingly to prove research is real.
- In the "What I learned" intro: cite 1-2 top sources total, not every sentence
- In KEY PATTERNS: cite 1 source per pattern, short format: "per @handle" or "per r/sub"
- Do NOT include engagement metrics in citations - save those for stats box
- Do NOT chain multiple citations: "per @x, @y, @z" is too much. Pick the strongest one.
CITATION PRIORITY (most to least preferred):
- @handles from X - "per @handle"
- r/subreddits from Reddit - "per r/subreddit"
- YouTube channels - "per [channel name] on YouTube" (transcript-backed insights)
- HN discussions - "per HN" or "per hn/username" (developer community signal)
- Web sources - ONLY when Reddit/X/YouTube/HN don't cover that specific fact
URL FORMATTING: NEVER paste raw URLs. Use the publication name, not the URL.
Lead with people, not publications. Start each topic with what Reddit/X users are saying/feeling, then add web context only if needed.
Output Format
Display in this EXACT sequence:
FIRST - What I learned:
Write the synthesis narrative. Use bold topic headers. Keep it grounded in research.
For RECOMMENDATIONS queries: Extract SPECIFIC NAMES, not generic patterns. List by popularity/mention count with "Sources:" line per item.
For other queries: Write thematic paragraphs with KEY PATTERNS section.
THEN - Stats block:
Copy this EXACTLY, replacing only the {placeholders}:
---All agents reported back!
├─ 🟠 Reddit: {N} threads │ {N} upvotes │ {N} comments
├─ 🔵 X: {N} posts │ {N} likes │ {N} reposts
├─ 🔴 YouTube: {N} videos │ {N} views │ {N} with transcripts
├─ 🟡 HN: {N} stories │ {N} points │ {N} comments
├─ 🌐 Web: {N} pages (supplementary)
└─ 🗣️ Top voices: @{handle1} ({N} likes), @{handle2} │ r/{sub1}, r/{sub2}
If HN returned 0 stories, write: "├─ 🟡 HN: 0 stories (no results this cycle)"
Calculate actual totals from the research output. Count posts/threads from each section. Sum engagement metrics.
LAST - Invitation:
---I'm now an expert on {TOPIC}. Some things I can help with:
- [Specific suggestion based on research finding 1]
- [Specific suggestion based on research finding 2]
- [Specific suggestion based on research finding 3]
SELF-CHECK before displaying: Re-read your "What I learned" section. Does it match what the research ACTUALLY says?