perf: optimize dedup, parallelize handle searches and enrichment
The dedup hot path recomputed normalize_text() 4 times per comparison and recomputed item_text() on every inner-loop iteration. Pre-computing n-gram sets and token sets into a _PreparedText cache cuts dedup time by 6x (2.16s to 0.39s on 300 unique items). Bird handle searches spawned one Node process per handle sequentially. Now uses ThreadPoolExecutor so N handles run concurrently. Same pattern applied to YouTube comment enrichment (was serial, Reddit was already parallel) and the retry-thin-sources phase in the pipeline. Clustering now pre-computes candidate text and uses prepared_similarity for the O(n^2) grouping and MMR representative selection loops. Minor: _is_wsl() cached with lru_cache, Bundle.add_items() uses extend() instead of list concatenation. End-to-end: 5.2s -> 3.7s (29% faster) on a typical 4-source query.
This commit is contained in:
@@ -7,6 +7,7 @@ Only uses Python stdlib — no external dependencies.
|
||||
"""
|
||||
|
||||
import configparser
|
||||
import functools
|
||||
import logging
|
||||
import platform
|
||||
import shutil
|
||||
@@ -18,8 +19,12 @@ from typing import Dict, List, Optional
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
|
||||
@functools.lru_cache(maxsize=1)
|
||||
def _is_wsl() -> bool:
|
||||
"""Detect if running under Windows Subsystem for Linux."""
|
||||
"""Detect if running under Windows Subsystem for Linux.
|
||||
|
||||
Cached after the first call since /proc/version doesn't change at runtime.
|
||||
"""
|
||||
try:
|
||||
return "microsoft" in Path("/proc/version").read_text().lower()
|
||||
except OSError:
|
||||
|
||||
Reference in New Issue
Block a user