diff --git a/CHANGELOG.md b/CHANGELOG.md index 8050048..57c2adf 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -9,7 +9,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 ### Added -- **Narrative lens for company / product / service topics.** Comparison tables gain a `Setting the narrative?` axis that judges whether each entity's community conversation is actually about what the entity *pitches* (its first-party positioning) or about something else — pricing, rivals, an incident, a ToS change. A new mandatory research step captures each entity's current stated positioning from first-party sources (homepage, docs, pricing) rather than from memory, and single-entity company runs get a `narrative-check` synthesis beat surfacing the same signal. The mismatch is the point: companies usually don't control their own conversation. The lens is scoped to entities with an identifiable first party (companies, products, services): people always get N/A — even founders whose companies would qualify — as do events, abstract concepts, and ownerless topics like Bitcoin, and verdicts require positioning fetched during the run, never from memory. +- **First-party positioning research + pitch-vs-pulse synthesis (company / product / service topics).** A new mandatory research step captures each entity's current stated positioning from first-party sources (homepage, docs, pricing) rather than from memory. The fetched pitch grounds `What it is` descriptions (entities described as they pitch themselves today), helps reject unrelated brand-name noise, and feeds an evidence-triggered prose beat: when the month's conversation directly supports a specific claim, cuts against one, or is squarely about the pitched ground, the synthesis says so anchored to the top thread — and stays silent when the pulse is orthogonal to the pitch, because a manufactured connection is worse than omission. Claims are tested at matched altitude (specific claims against specific threads; broad taglines are never graded against individual items), and statements stay windowed to the 30 days — no trend verdicts. Scoped to entities with an identifiable first party: people are always excluded (even founders whose companies qualify), as are events, abstract concepts, and ownerless topics like Bitcoin; the beat requires positioning fetched during the run, never from memory. ### Fixed diff --git a/skills/last30days/SKILL.md b/skills/last30days/SKILL.md index 34aa650..e7e1724 100644 --- a/skills/last30days/SKILL.md +++ b/skills/last30days/SKILL.md @@ -641,11 +641,11 @@ Topic A (the main topic, first in the vs-string) uses outer `--x-handle`, `--x-r **Then do WebSearch supplements** for: `{TOPIC_A} vs {TOPIC_B} comparison {YEAR}` and `{TOPIC_A} vs {TOPIC_B} which is better` — these catch rivalry articles that per-entity passes might not surface. -**Fill the `Setting the narrative?` row using `RESOLVED_POSITIONING` + the community evidence.** Run Step 0.55 item 6 (first-party positioning) per entity - it gives you each entity's CURRENT stated pitch, fetched fresh, not from memory. Then judge: is the community conversation actually ABOUT what the entity pitches, or about something else (pricing, rivals, an incident, a ToS change)? Start the cell with **Yes / Partly / No / Unclear**, then name the topic the community is really on, anchored to a real item with its engagement (e.g. "No - pitches uptime, but the top thread is friendly-fraud, 323pt HN"). The mismatch is the signal: companies usually don't control their own conversation. Use **Unclear** when the evidence is thin or polluted with unrelated brand-name matches (e.g. someone's `.vercel.app` app, an exam "calc") - do NOT infer a verdict from vibes. Write N/A for entities with no public pitch: people (ALWAYS, even founders/CEOs whose companies would qualify - "Garry Tan vs Sam Altman" gets N/A across the row), events, abstract concepts, and ownerless topics (Bitcoin - a foundation or fan site is not an authoritative first party). And only write a verdict when positioning was actually FETCHED this run: if item 6 could not run (e.g. no WebSearch), write Unclear - never supply the pitch from memory. +**Use `RESOLVED_POSITIONING` per entity (Step 0.55 item 6) in two ways.** First, ground each entity's `What it is` cell in its CURRENT fetched pitch - describe the entity as it pitches itself today, never from memory. Second, if an entity's month of evidence directly bears on its pitch - SUPPORTS a specific claim, CUTS AGAINST one, or the conversation is squarely ABOUT the pitched ground - say so in PROSE inside that entity's narrative section, anchored to the real item with its engagement. When the pulse is orthogonal to the pitch (on-entity but about something the pitch doesn't speak to), say NOTHING about the pitch: omission is the correct output, and a manufactured connection is worse than silence. Match altitude: test SPECIFIC claims ("zero-config", "fastest", an uptime number) against specific threads; never grade a broad tagline ("financial infrastructure") against an individual thread - it is too broad to hit or miss. **Skip the normal Step 1 below** - go directly to the comparison synthesis format (see "If QUERY_TYPE = COMPARISON" in the synthesis section). -**COMPARISON TABLE SCAFFOLD (engine-emitted, pass through verbatim):** For comparison topics, the engine's compact output includes a `## Head-to-Head` block with an empty markdown table (columns = entities, rows = axes like "What it is", "Setting the narrative?", "Best for"). Your synthesis MUST include this block verbatim with filled cells, positioned between the narrative and the emoji-tree footer. Keep each cell to 5-15 words. Use ' - ' (hyphen with spaces) not em-dashes inside cells. +**COMPARISON TABLE SCAFFOLD (engine-emitted, pass through verbatim):** For comparison topics, the engine's compact output includes a `## Head-to-Head` block with an empty markdown table (columns = entities, rows = axes like "What it is", "Philosophy", "Best for"). Your synthesis MUST include this block verbatim with filled cells, positioned between the narrative and the emoji-tree footer. Keep each cell to 5-15 words. Use ' - ' (hyphen with spaces) not em-dashes inside cells. ### Competitor mode (`--competitors`) @@ -750,7 +750,7 @@ Store as `RESOLVED_IG_CREATORS`. Store as `RESOLVED_YT_QUERIES`. -**6. First-party positioning** - **MANDATORY when WebSearch is available, for company / product / service topics.** If the topic (or, in a vs-run, an entity) is a company, product, or service with a public presence, fetch its CURRENT stated positioning. Do **NOT** rely on memory - homepages and positioning go stale as companies rewrite copy and pivot, and a stale claim produces a false gap. Anchor on first-party sources: the homepage tagline, docs, pricing, or a "compare/why-us" page. Fold this into the per-entity passes above where you can (e.g. add `official site` to a query); otherwise run one focused search per entity (`{TOPIC} official site`, `{TOPIC} pricing`). Capture the one-line value prop and any explicit claims ("zero-config", "fastest", "open source"). Store as `RESOLVED_POSITIONING`. This is what the entity *pitches*; the engine's community data is what people *actually talk about*. Whether those two line up feeds the `Setting the narrative?` synthesis (the conversation is often on a different topic than the pitch - that mismatch is the signal). Skip (and omit `RESOLVED_POSITIONING`) for people, events, abstract concepts, and ownerless topics - they make no comparable public claim. The test is an identifiable first party with a fetchable pitch, and people NEVER pass it - not even founders/creators whose companies would qualify. The lens can apply to MrBeast (a company) but never to Jimmy Donaldson (a person); a person-vs-person run ("Garry Tan vs Sam Altman") gets no positioning research at all. Ownerless topics fail the same test: Bitcoin has no authoritative first party, and a foundation or fan site does not count. +**6. First-party positioning** - **MANDATORY when WebSearch is available, for company / product / service topics.** If the topic (or, in a vs-run, an entity) is a company, product, or service with a public presence, fetch its CURRENT stated positioning. Do **NOT** rely on memory - homepages and positioning go stale as companies rewrite copy and pivot, and a stale claim produces a false gap. Anchor on first-party sources: the homepage tagline, docs, pricing, or a "compare/why-us" page. Fold this into the per-entity passes above where you can (e.g. add `official site` to a query); otherwise run one focused search per entity (`{TOPIC} official site`, `{TOPIC} pricing`). Capture the one-line value prop and any explicit claims ("zero-config", "fastest", "open source"). Store as `RESOLVED_POSITIONING`. This is what the entity *pitches*; the engine's community data is what people *actually talk about*. Use it three ways: ground `What it is` descriptions (describe the entity as it pitches itself TODAY, not as remembered), help reject unrelated brand-name noise (knowing what the entity is makes off-brand matches obvious), and feed the pitch-vs-pulse synthesis beat - a PROSE note that fires only when the month's evidence directly supports, cuts against, or is squarely about the pitch (see the synthesis section; orthogonal evidence gets silence, not a verdict). Skip (and omit `RESOLVED_POSITIONING`) for people, events, abstract concepts, and ownerless topics - they make no comparable public claim. The test is an identifiable first party with a fetchable pitch, and people NEVER pass it - not even founders/creators whose companies would qualify. The lens can apply to MrBeast (a company) but never to Jimmy Donaldson (a person); a person-vs-person run ("Garry Tan vs Sam Altman") gets no positioning research at all. Ownerless topics fail the same test: Bitcoin has no authoritative first party, and a foundation or fan site does not count. **Concrete examples:** @@ -1295,7 +1295,6 @@ Voice contract LAWs 1, 3, 5 apply to comparisons unchanged (no `Sources:` block, | Dimension | {Entity 1} | {Entity 2} | {Entity 3} | |---|---|---|---| | What it is | ... | ... | ... | -| Setting the narrative? | ... | ... | ... | | GitHub stars | ... | ... | ... | | Philosophy | ... | ... | ... | | Skills | ... | ... | ... | @@ -1305,7 +1304,7 @@ Voice contract LAWs 1, 3, 5 apply to comparisons unchanged (no `Sources:` block, | Best for | ... | ... | ... | | Install | ... | ... | ... | -(Engine emits this scaffold; fill the cells with 5-15 words each. If an axis does not apply to the topic class, write "N/A" or a topic-appropriate substitute rather than inventing data. For the `Setting the narrative?` row, judge whether each entity's community conversation is about what the entity itself pitches: start the cell with **Yes / Partly / No / Unclear**, then name the topic the community is ACTUALLY on, anchored to a real item with engagement, e.g. "No - pitches uptime, but the top thread is friendly-fraud (323pt HN)". The mismatch is the signal - companies usually don't control their own conversation. Use "Unclear" when evidence is thin or polluted with unrelated brand-name matches; write N/A for entities with no public pitch - people (even famous founders), events, abstract concepts, and ownerless topics like Bitcoin. Verdicts require positioning fetched THIS run; if it wasn't fetched, write Unclear rather than pitching from memory. Do NOT infer a verdict from vibes.) +(Engine emits this scaffold; fill the cells with 5-15 words each. If an axis does not apply to the topic class, write "N/A" or a topic-appropriate substitute rather than inventing data. Ground the `What it is` row in `RESOLVED_POSITIONING` when captured - each entity described as it pitches itself today, fetched this run, never from memory.) ## The Bottom Line @@ -1447,7 +1446,7 @@ At render time the `@handle`, `r/sub`, and publication-name placeholders become Headlines should be specific and newsy ("BULLY dropped and it's dominating", "Europe is banning him one country at a time"), not generic ("Album release", "Tour updates"). -**Narrative-check beat (company / product / service topics).** If the topic is a company, product, or service and you captured `RESOLVED_POSITIONING` in Step 0.55, work in ONE bold-lead-in paragraph on whether the entity is setting its own narrative - i.e. is the community actually talking about what it pitches, or about something else (pricing, rivals, an incident, a ToS change)? Anchor the verdict to the real top-discussed item with its engagement - e.g. `**Vercel is losing the narrative to its pricing** - it still sells "the AI Cloud" and zero-overhead speed, but the loudest community thread this month is about cost and a June 1 ToS change, not performance`. When the conversation DOES track the pitch, say so (that is the "Yes" case, equally worth stating). Keep it a normal newsy bold-lead-in paragraph with a specific headline (per the rule above), NOT a new `##` section (LAW 4 still holds). Skip it silently for people (always - the beat can cover MrBeast the company, never Jimmy Donaldson the person), events, abstract concepts, and ownerless topics (Bitcoin), or when the evidence is too thin or noise-polluted to tell - do not manufacture a verdict. +**Pitch-vs-pulse beat (company / product / service topics).** If you captured `RESOLVED_POSITIONING` in Step 0.55 AND the month's evidence directly bears on it, work in ONE bold-lead-in paragraph saying how. Three cases qualify: the pulse SUPPORTS a specific claim (e.g. `**"Zero-config" is holding up** - this month's top deploy thread is devs praising the no-setup flow, 800 upvotes`), CUTS AGAINST one (e.g. `**Stripe's fraud-fighting pitch took a direct hit** - the loudest thread this month argues it is friendly to "friendly fraud", 323pt HN`), or the conversation is squarely ABOUT the pitched ground. Always anchor to the real top item with its engagement, and keep claims windowed - "this month's conversation" - never trend verbs like "losing the narrative" that one 30-day window cannot support. If the month's conversation is orthogonal to the pitch - on-entity but about something the pitch doesn't speak to - write NOTHING about the pitch: omission is the correct output, and a manufactured connection is worse than silence. Match altitude: test SPECIFIC claims ("zero-config", "fastest", an uptime number) against specific threads; never grade a broad tagline against an individual thread. Keep it a normal newsy bold-lead-in paragraph, NOT a new `##` section (LAW 4 still holds). Skip silently for people (always - the beat can cover MrBeast the company, never Jimmy Donaldson the person), events, abstract concepts, and ownerless topics (Bitcoin), and whenever positioning was not actually fetched this run - never supply a pitch from memory. **THEN - Quality Nudge (if present in the output):** diff --git a/skills/last30days/scripts/lib/render.py b/skills/last30days/scripts/lib/render.py index 845155a..dc61dbb 100644 --- a/skills/last30days/scripts/lib/render.py +++ b/skills/last30days/scripts/lib/render.py @@ -531,12 +531,10 @@ def _render_comparison_scaffold(topic: str) -> list[str]: Returns empty list if topic is not a comparison query. When present, the block is bracketed so the synthesizer can detect it and pass through. - Axes are the 9 from the April 9 launch-video exemplar (suited to AI-tool - comparisons) plus "Setting the narrative?", which asks whether each entity's - community conversation is about what the entity itself pitches, or about - something else (pricing, rivals, incidents). For comparisons where an axis - does not apply, the synthesizer writes N/A or a topic-appropriate substitute - in that row. + Axes match the April 9 launch-video exemplar (9 axes suited to AI-tool + comparisons). For non-AI-tool comparisons, the synthesizer writes N/A + or topic-appropriate substitutes in irrelevant rows. The "What it is" row + grounds in first-party positioning fetched during the run when available. """ entities = _parse_comparison_entities(topic) if not entities: @@ -546,11 +544,10 @@ def _render_comparison_scaffold(topic: str) -> list[str]: header = "| Dimension | " + " | ".join(entities) + " |" # Separator row matching column count separator = "|" + "|".join(["---"] * (len(entities) + 1)) + "|" - # 9 axes from the April 9 exemplar plus "Setting the narrative?" (pitch vs. - # what the community actually talks about). See the function docstring. + # 9 axes from the April 9 exemplar. Model fills with topic-appropriate + # content; irrelevant axes get "N/A" rather than invented data. axes = [ "What it is", - "Setting the narrative?", "GitHub stars", "Philosophy", "Skills", @@ -562,20 +559,11 @@ def _render_comparison_scaffold(topic: str) -> list[str]: ] body = [f"| {axis} | " + " | ".join([" "] * len(entities)) + " |" for axis in axes] - # Generic fill rules plus guidance for "Setting the narrative?" - the one - # axis that needs a judgement, not a lookup. fill_instructions = ( "Fill each cell based on the research above. Keep cells short (5-15 words). " "Use ' - ' (hyphen with spaces) not em-dashes. Write N/A for axes that do not apply to this topic class. " - "For the \"Setting the narrative?\" row, judge whether each entity's community conversation is about " - "what the entity itself pitches: start the cell with Yes / Partly / No / Unclear, then name the topic " - "the community is ACTUALLY on, anchored to a real item (e.g. \"No - pitches uptime, but the top thread " - "is friendly-fraud (323pt HN)\"). Use Unclear when evidence is thin or polluted with unrelated " - "brand-name matches; do NOT infer a verdict from vibes. Write N/A for entities with no public pitch: " - "people (always, even famous founders - the lens applies to companies and products, never persons), " - "events, abstract concepts, and ownerless topics with no authoritative first party (e.g. Bitcoin). " - "Only fill a verdict when the entity's first-party positioning was actually fetched during this " - "run's research; if it was not, write Unclear - never supply the pitch from memory. " + "Ground the \"What it is\" row in first-party positioning fetched during this run's research when " + "available - describe each entity as it pitches itself today, never from memory. " "This scaffold matches the April 9 launch-video exemplar shape." ) diff --git a/tests/test_render_comparison_multi.py b/tests/test_render_comparison_multi.py index ee44f07..7b61275 100644 --- a/tests/test_render_comparison_multi.py +++ b/tests/test_render_comparison_multi.py @@ -101,11 +101,11 @@ class RenderComparisonMultiTests(unittest.TestCase): self.assertIn("## xAI", rendered) # Scaffold table header has a column per entity self.assertIn("| Dimension | OpenAI | Anthropic | xAI |", rendered) - # The narrative-lens axis is emitted as a scaffold row - self.assertIn("| Setting the narrative? |", rendered) - # Fill instructions carry the scope + artifact gate (no verdicts for - # people/ownerless topics, no pitch from memory) - self.assertIn("never supply the pitch from memory", rendered) + # No verdict row: the pitch-vs-pulse signal ships as synthesis prose, + # not a table axis (early drafts emitted a "Setting the narrative?" row) + self.assertNotIn("Setting the narrative?", rendered) + # "What it is" grounds in positioning fetched this run, never memory + self.assertIn("never from memory", rendered) # Envelope scaffolding present self.assertIn("EVIDENCE FOR SYNTHESIS", rendered) self.assertIn("END OF last30days CANONICAL OUTPUT", rendered)