feat(skill): replace narrative verdict row with evidence-triggered prose beat

Review showed the "Setting the narrative?" verdict compared across
abstraction levels: a homepage tagline is deliberately broad
("financial infrastructure" covers a chargebacks thread), so
tagline-vs-thread alignment verdicts are unfalsifiable and carry no
information. The signal now ships as PROSE in the entity's narrative
section, fires only when the month's evidence directly bears on the
pitch (supports a specific claim, cuts against one, or is squarely
about the pitched ground), and stays SILENT when the pulse is
orthogonal - omission over a manufactured connection. Claims are
tested at matched altitude (specific claim vs specific thread) and
stay windowed (no trend verbs one 30-day window can't support). The
positioning fetch step survives unchanged and now also grounds the
"What it is" row and brand-noise rejection. All scope gating (people
never, ownerless topics excluded, no pitch from memory) carries over.
This commit is contained in:
Trevin Chow
2026-06-09 17:14:31 -07:00
parent 57860aff1c
commit c1ca1a4e9d
4 changed files with 19 additions and 32 deletions
+1 -1
View File
@@ -9,7 +9,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
### Added
- **Narrative lens for company / product / service topics.** Comparison tables gain a `Setting the narrative?` axis that judges whether each entity's community conversation is actually about what the entity *pitches* (its first-party positioning) or about something else — pricing, rivals, an incident, a ToS change. A new mandatory research step captures each entity's current stated positioning from first-party sources (homepage, docs, pricing) rather than from memory, and single-entity company runs get a `narrative-check` synthesis beat surfacing the same signal. The mismatch is the point: companies usually don't control their own conversation. The lens is scoped to entities with an identifiable first party (companies, products, services): people always get N/A — even founders whose companies would qualify as do events, abstract concepts, and ownerless topics like Bitcoin, and verdicts require positioning fetched during the run, never from memory.
- **First-party positioning research + pitch-vs-pulse synthesis (company / product / service topics).** A new mandatory research step captures each entity's current stated positioning from first-party sources (homepage, docs, pricing) rather than from memory. The fetched pitch grounds `What it is` descriptions (entities described as they pitch themselves today), helps reject unrelated brand-name noise, and feeds an evidence-triggered prose beat: when the month's conversation directly supports a specific claim, cuts against one, or is squarely about the pitched ground, the synthesis says so anchored to the top thread — and stays silent when the pulse is orthogonal to the pitch, because a manufactured connection is worse than omission. Claims are tested at matched altitude (specific claims against specific threads; broad taglines are never graded against individual items), and statements stay windowed to the 30 days — no trend verdicts. Scoped to entities with an identifiable first party: people are always excluded (even founders whose companies qualify), as are events, abstract concepts, and ownerless topics like Bitcoin; the beat requires positioning fetched during the run, never from memory.
### Fixed
+5 -6
View File
@@ -641,11 +641,11 @@ Topic A (the main topic, first in the vs-string) uses outer `--x-handle`, `--x-r
**Then do WebSearch supplements** for: `{TOPIC_A} vs {TOPIC_B} comparison {YEAR}` and `{TOPIC_A} vs {TOPIC_B} which is better` — these catch rivalry articles that per-entity passes might not surface.
**Fill the `Setting the narrative?` row using `RESOLVED_POSITIONING` + the community evidence.** Run Step 0.55 item 6 (first-party positioning) per entity - it gives you each entity's CURRENT stated pitch, fetched fresh, not from memory. Then judge: is the community conversation actually ABOUT what the entity pitches, or about something else (pricing, rivals, an incident, a ToS change)? Start the cell with **Yes / Partly / No / Unclear**, then name the topic the community is really on, anchored to a real item with its engagement (e.g. "No - pitches uptime, but the top thread is friendly-fraud, 323pt HN"). The mismatch is the signal: companies usually don't control their own conversation. Use **Unclear** when the evidence is thin or polluted with unrelated brand-name matches (e.g. someone's `.vercel.app` app, an exam "calc") - do NOT infer a verdict from vibes. Write N/A for entities with no public pitch: people (ALWAYS, even founders/CEOs whose companies would qualify - "Garry Tan vs Sam Altman" gets N/A across the row), events, abstract concepts, and ownerless topics (Bitcoin - a foundation or fan site is not an authoritative first party). And only write a verdict when positioning was actually FETCHED this run: if item 6 could not run (e.g. no WebSearch), write Unclear - never supply the pitch from memory.
**Use `RESOLVED_POSITIONING` per entity (Step 0.55 item 6) in two ways.** First, ground each entity's `What it is` cell in its CURRENT fetched pitch - describe the entity as it pitches itself today, never from memory. Second, if an entity's month of evidence directly bears on its pitch - SUPPORTS a specific claim, CUTS AGAINST one, or the conversation is squarely ABOUT the pitched ground - say so in PROSE inside that entity's narrative section, anchored to the real item with its engagement. When the pulse is orthogonal to the pitch (on-entity but about something the pitch doesn't speak to), say NOTHING about the pitch: omission is the correct output, and a manufactured connection is worse than silence. Match altitude: test SPECIFIC claims ("zero-config", "fastest", an uptime number) against specific threads; never grade a broad tagline ("financial infrastructure") against an individual thread - it is too broad to hit or miss.
**Skip the normal Step 1 below** - go directly to the comparison synthesis format (see "If QUERY_TYPE = COMPARISON" in the synthesis section).
**COMPARISON TABLE SCAFFOLD (engine-emitted, pass through verbatim):** For comparison topics, the engine's compact output includes a `## Head-to-Head` block with an empty markdown table (columns = entities, rows = axes like "What it is", "Setting the narrative?", "Best for"). Your synthesis MUST include this block verbatim with filled cells, positioned between the narrative and the emoji-tree footer. Keep each cell to 5-15 words. Use ' - ' (hyphen with spaces) not em-dashes inside cells.
**COMPARISON TABLE SCAFFOLD (engine-emitted, pass through verbatim):** For comparison topics, the engine's compact output includes a `## Head-to-Head` block with an empty markdown table (columns = entities, rows = axes like "What it is", "Philosophy", "Best for"). Your synthesis MUST include this block verbatim with filled cells, positioned between the narrative and the emoji-tree footer. Keep each cell to 5-15 words. Use ' - ' (hyphen with spaces) not em-dashes inside cells.
### Competitor mode (`--competitors`)
@@ -750,7 +750,7 @@ Store as `RESOLVED_IG_CREATORS`.
Store as `RESOLVED_YT_QUERIES`.
**6. First-party positioning** - **MANDATORY when WebSearch is available, for company / product / service topics.** If the topic (or, in a vs-run, an entity) is a company, product, or service with a public presence, fetch its CURRENT stated positioning. Do **NOT** rely on memory - homepages and positioning go stale as companies rewrite copy and pivot, and a stale claim produces a false gap. Anchor on first-party sources: the homepage tagline, docs, pricing, or a "compare/why-us" page. Fold this into the per-entity passes above where you can (e.g. add `official site` to a query); otherwise run one focused search per entity (`{TOPIC} official site`, `{TOPIC} pricing`). Capture the one-line value prop and any explicit claims ("zero-config", "fastest", "open source"). Store as `RESOLVED_POSITIONING`. This is what the entity *pitches*; the engine's community data is what people *actually talk about*. Whether those two line up feeds the `Setting the narrative?` synthesis (the conversation is often on a different topic than the pitch - that mismatch is the signal). Skip (and omit `RESOLVED_POSITIONING`) for people, events, abstract concepts, and ownerless topics - they make no comparable public claim. The test is an identifiable first party with a fetchable pitch, and people NEVER pass it - not even founders/creators whose companies would qualify. The lens can apply to MrBeast (a company) but never to Jimmy Donaldson (a person); a person-vs-person run ("Garry Tan vs Sam Altman") gets no positioning research at all. Ownerless topics fail the same test: Bitcoin has no authoritative first party, and a foundation or fan site does not count.
**6. First-party positioning** - **MANDATORY when WebSearch is available, for company / product / service topics.** If the topic (or, in a vs-run, an entity) is a company, product, or service with a public presence, fetch its CURRENT stated positioning. Do **NOT** rely on memory - homepages and positioning go stale as companies rewrite copy and pivot, and a stale claim produces a false gap. Anchor on first-party sources: the homepage tagline, docs, pricing, or a "compare/why-us" page. Fold this into the per-entity passes above where you can (e.g. add `official site` to a query); otherwise run one focused search per entity (`{TOPIC} official site`, `{TOPIC} pricing`). Capture the one-line value prop and any explicit claims ("zero-config", "fastest", "open source"). Store as `RESOLVED_POSITIONING`. This is what the entity *pitches*; the engine's community data is what people *actually talk about*. Use it three ways: ground `What it is` descriptions (describe the entity as it pitches itself TODAY, not as remembered), help reject unrelated brand-name noise (knowing what the entity is makes off-brand matches obvious), and feed the pitch-vs-pulse synthesis beat - a PROSE note that fires only when the month's evidence directly supports, cuts against, or is squarely about the pitch (see the synthesis section; orthogonal evidence gets silence, not a verdict). Skip (and omit `RESOLVED_POSITIONING`) for people, events, abstract concepts, and ownerless topics - they make no comparable public claim. The test is an identifiable first party with a fetchable pitch, and people NEVER pass it - not even founders/creators whose companies would qualify. The lens can apply to MrBeast (a company) but never to Jimmy Donaldson (a person); a person-vs-person run ("Garry Tan vs Sam Altman") gets no positioning research at all. Ownerless topics fail the same test: Bitcoin has no authoritative first party, and a foundation or fan site does not count.
**Concrete examples:**
@@ -1295,7 +1295,6 @@ Voice contract LAWs 1, 3, 5 apply to comparisons unchanged (no `Sources:` block,
| Dimension | {Entity 1} | {Entity 2} | {Entity 3} |
|---|---|---|---|
| What it is | ... | ... | ... |
| Setting the narrative? | ... | ... | ... |
| GitHub stars | ... | ... | ... |
| Philosophy | ... | ... | ... |
| Skills | ... | ... | ... |
@@ -1305,7 +1304,7 @@ Voice contract LAWs 1, 3, 5 apply to comparisons unchanged (no `Sources:` block,
| Best for | ... | ... | ... |
| Install | ... | ... | ... |
(Engine emits this scaffold; fill the cells with 5-15 words each. If an axis does not apply to the topic class, write "N/A" or a topic-appropriate substitute rather than inventing data. For the `Setting the narrative?` row, judge whether each entity's community conversation is about what the entity itself pitches: start the cell with **Yes / Partly / No / Unclear**, then name the topic the community is ACTUALLY on, anchored to a real item with engagement, e.g. "No - pitches uptime, but the top thread is friendly-fraud (323pt HN)". The mismatch is the signal - companies usually don't control their own conversation. Use "Unclear" when evidence is thin or polluted with unrelated brand-name matches; write N/A for entities with no public pitch - people (even famous founders), events, abstract concepts, and ownerless topics like Bitcoin. Verdicts require positioning fetched THIS run; if it wasn't fetched, write Unclear rather than pitching from memory. Do NOT infer a verdict from vibes.)
(Engine emits this scaffold; fill the cells with 5-15 words each. If an axis does not apply to the topic class, write "N/A" or a topic-appropriate substitute rather than inventing data. Ground the `What it is` row in `RESOLVED_POSITIONING` when captured - each entity described as it pitches itself today, fetched this run, never from memory.)
## The Bottom Line
@@ -1447,7 +1446,7 @@ At render time the `@handle`, `r/sub`, and publication-name placeholders become
Headlines should be specific and newsy ("BULLY dropped and it's dominating", "Europe is banning him one country at a time"), not generic ("Album release", "Tour updates").
**Narrative-check beat (company / product / service topics).** If the topic is a company, product, or service and you captured `RESOLVED_POSITIONING` in Step 0.55, work in ONE bold-lead-in paragraph on whether the entity is setting its own narrative - i.e. is the community actually talking about what it pitches, or about something else (pricing, rivals, an incident, a ToS change)? Anchor the verdict to the real top-discussed item with its engagement - e.g. `**Vercel is losing the narrative to its pricing** - it still sells "the AI Cloud" and zero-overhead speed, but the loudest community thread this month is about cost and a June 1 ToS change, not performance`. When the conversation DOES track the pitch, say so (that is the "Yes" case, equally worth stating). Keep it a normal newsy bold-lead-in paragraph with a specific headline (per the rule above), NOT a new `##` section (LAW 4 still holds). Skip it silently for people (always - the beat can cover MrBeast the company, never Jimmy Donaldson the person), events, abstract concepts, and ownerless topics (Bitcoin), or when the evidence is too thin or noise-polluted to tell - do not manufacture a verdict.
**Pitch-vs-pulse beat (company / product / service topics).** If you captured `RESOLVED_POSITIONING` in Step 0.55 AND the month's evidence directly bears on it, work in ONE bold-lead-in paragraph saying how. Three cases qualify: the pulse SUPPORTS a specific claim (e.g. `**"Zero-config" is holding up** - this month's top deploy thread is devs praising the no-setup flow, 800 upvotes`), CUTS AGAINST one (e.g. `**Stripe's fraud-fighting pitch took a direct hit** - the loudest thread this month argues it is friendly to "friendly fraud", 323pt HN`), or the conversation is squarely ABOUT the pitched ground. Always anchor to the real top item with its engagement, and keep claims windowed - "this month's conversation" - never trend verbs like "losing the narrative" that one 30-day window cannot support. If the month's conversation is orthogonal to the pitch - on-entity but about something the pitch doesn't speak to - write NOTHING about the pitch: omission is the correct output, and a manufactured connection is worse than silence. Match altitude: test SPECIFIC claims ("zero-config", "fastest", an uptime number) against specific threads; never grade a broad tagline against an individual thread. Keep it a normal newsy bold-lead-in paragraph, NOT a new `##` section (LAW 4 still holds). Skip silently for people (always - the beat can cover MrBeast the company, never Jimmy Donaldson the person), events, abstract concepts, and ownerless topics (Bitcoin), and whenever positioning was not actually fetched this run - never supply a pitch from memory.
**THEN - Quality Nudge (if present in the output):**
+8 -20
View File
@@ -531,12 +531,10 @@ def _render_comparison_scaffold(topic: str) -> list[str]:
Returns empty list if topic is not a comparison query. When present,
the block is bracketed so the synthesizer can detect it and pass through.
Axes are the 9 from the April 9 launch-video exemplar (suited to AI-tool
comparisons) plus "Setting the narrative?", which asks whether each entity's
community conversation is about what the entity itself pitches, or about
something else (pricing, rivals, incidents). For comparisons where an axis
does not apply, the synthesizer writes N/A or a topic-appropriate substitute
in that row.
Axes match the April 9 launch-video exemplar (9 axes suited to AI-tool
comparisons). For non-AI-tool comparisons, the synthesizer writes N/A
or topic-appropriate substitutes in irrelevant rows. The "What it is" row
grounds in first-party positioning fetched during the run when available.
"""
entities = _parse_comparison_entities(topic)
if not entities:
@@ -546,11 +544,10 @@ def _render_comparison_scaffold(topic: str) -> list[str]:
header = "| Dimension | " + " | ".join(entities) + " |"
# Separator row matching column count
separator = "|" + "|".join(["---"] * (len(entities) + 1)) + "|"
# 9 axes from the April 9 exemplar plus "Setting the narrative?" (pitch vs.
# what the community actually talks about). See the function docstring.
# 9 axes from the April 9 exemplar. Model fills with topic-appropriate
# content; irrelevant axes get "N/A" rather than invented data.
axes = [
"What it is",
"Setting the narrative?",
"GitHub stars",
"Philosophy",
"Skills",
@@ -562,20 +559,11 @@ def _render_comparison_scaffold(topic: str) -> list[str]:
]
body = [f"| {axis} | " + " | ".join([" "] * len(entities)) + " |" for axis in axes]
# Generic fill rules plus guidance for "Setting the narrative?" - the one
# axis that needs a judgement, not a lookup.
fill_instructions = (
"Fill each cell based on the research above. Keep cells short (5-15 words). "
"Use ' - ' (hyphen with spaces) not em-dashes. Write N/A for axes that do not apply to this topic class. "
"For the \"Setting the narrative?\" row, judge whether each entity's community conversation is about "
"what the entity itself pitches: start the cell with Yes / Partly / No / Unclear, then name the topic "
"the community is ACTUALLY on, anchored to a real item (e.g. \"No - pitches uptime, but the top thread "
"is friendly-fraud (323pt HN)\"). Use Unclear when evidence is thin or polluted with unrelated "
"brand-name matches; do NOT infer a verdict from vibes. Write N/A for entities with no public pitch: "
"people (always, even famous founders - the lens applies to companies and products, never persons), "
"events, abstract concepts, and ownerless topics with no authoritative first party (e.g. Bitcoin). "
"Only fill a verdict when the entity's first-party positioning was actually fetched during this "
"run's research; if it was not, write Unclear - never supply the pitch from memory. "
"Ground the \"What it is\" row in first-party positioning fetched during this run's research when "
"available - describe each entity as it pitches itself today, never from memory. "
"This scaffold matches the April 9 launch-video exemplar shape."
)
+5 -5
View File
@@ -101,11 +101,11 @@ class RenderComparisonMultiTests(unittest.TestCase):
self.assertIn("## xAI", rendered)
# Scaffold table header has a column per entity
self.assertIn("| Dimension | OpenAI | Anthropic | xAI |", rendered)
# The narrative-lens axis is emitted as a scaffold row
self.assertIn("| Setting the narrative? |", rendered)
# Fill instructions carry the scope + artifact gate (no verdicts for
# people/ownerless topics, no pitch from memory)
self.assertIn("never supply the pitch from memory", rendered)
# No verdict row: the pitch-vs-pulse signal ships as synthesis prose,
# not a table axis (early drafts emitted a "Setting the narrative?" row)
self.assertNotIn("Setting the narrative?", rendered)
# "What it is" grounds in positioning fetched this run, never memory
self.assertIn("never from memory", rendered)
# Envelope scaffolding present
self.assertIn("EVIDENCE FOR SYNTHESIS", rendered)
self.assertIn("END OF last30days CANONICAL OUTPUT", rendered)