perf: cache PreparedQuery per stream, skip double-normalize in dedupe (#282)
Scoring hot path (_normalize_score_dedupe) re-tokenized the same ranking_query ~240x per stream: once per item for local_relevance, plus ~5x per item across snippet windows. Query tokens are immutable within a stream, so compute them once as relevance.PreparedQuery and thread through signals.annotate_stream and snippet.extract_best_snippet. dedupe._PreparedText called normalize_text twice: once in __init__ and again via get_ngrams. Factor out _ngrams_of_normalized so the prepared path skips the redundant pass while get_ngrams keeps its public contract. Behavior unchanged.
This commit is contained in:
@@ -26,7 +26,7 @@ def _windows(words: list[str], size: int, overlap: int) -> list[str]:
|
||||
|
||||
def extract_best_snippet(
|
||||
item: schema.SourceItem,
|
||||
ranking_query: str,
|
||||
ranking_query: "str | relevance.PreparedQuery",
|
||||
max_words: int = 120,
|
||||
) -> str:
|
||||
"""Prefer existing snippets, else extract the best matching evidence window."""
|
||||
@@ -43,8 +43,9 @@ def extract_best_snippet(
|
||||
if not candidates:
|
||||
return _truncate_words(body, max_words)
|
||||
|
||||
prepared_query = ranking_query if isinstance(ranking_query, relevance.PreparedQuery) else relevance.PreparedQuery(ranking_query)
|
||||
best = max(
|
||||
candidates,
|
||||
key=lambda candidate: relevance.token_overlap_relevance(ranking_query, candidate),
|
||||
key=lambda candidate: relevance.token_overlap_relevance(prepared_query, candidate),
|
||||
)
|
||||
return _truncate_words(best, max_words)
|
||||
|
||||
Reference in New Issue
Block a user