perf: cache PreparedQuery per stream, skip double-normalize in dedupe (#282)

Scoring hot path (_normalize_score_dedupe) re-tokenized the same
ranking_query ~240x per stream: once per item for local_relevance,
plus ~5x per item across snippet windows. Query tokens are immutable
within a stream, so compute them once as relevance.PreparedQuery and
thread through signals.annotate_stream and snippet.extract_best_snippet.

dedupe._PreparedText called normalize_text twice: once in __init__ and
again via get_ngrams. Factor out _ngrams_of_normalized so the prepared
path skips the redundant pass while get_ngrams keeps its public contract.

Behavior unchanged.
This commit is contained in:
Ilia Alshanetsky
2026-04-25 17:16:57 -04:00
committed by GitHub
parent 2c2755b49c
commit e6b89f2644
5 changed files with 48 additions and 18 deletions
+4 -2
View File
@@ -30,6 +30,7 @@ from . import (
query,
reddit,
reddit_public,
relevance,
rerank,
schema,
signals,
@@ -500,11 +501,12 @@ def _normalize_score_dedupe(
source, raw_items, from_date, to_date,
freshness_mode=freshness_mode,
)
normalized = signals.annotate_stream(normalized, ranking_query, freshness_mode)
prepared_query = relevance.PreparedQuery(ranking_query)
normalized = signals.annotate_stream(normalized, prepared_query, freshness_mode)
normalized = signals.prune_low_relevance(normalized)
normalized = dedupe.dedupe_items(normalized)
for item in normalized:
item.snippet = snippet.extract_best_snippet(item, ranking_query)
item.snippet = snippet.extract_best_snippet(item, prepared_query)
return normalized