Address review feedback: deduplicate query_type, clean unused imports, fix defaults

- Remove duplicate detect_query_type from query.py (divergent 5-type version);
  canonical 7-type version lives in query_type.py
- Fix reddit.py import to use query_type.detect_query_type
- Clean unused STOPWORDS/SYNONYMS/tokenize imports from youtube_yt, instagram,
  tiktok, scrapecreators_x, bird_x after relevance consolidation
- Fix _relevance_filter default from 0.7 to 0.0 (items without relevance
  should not silently pass the filter)
- Remove --dateafter from yt-dlp (returns 0 results for evergreen topics)
- Remove restrictSearchableAttributes from HN search (misses Ask/Show HN)
- Lower HN points filter from >5 to >2 (avoids filtering niche posts)
- Add error logging to select_openai_model HTTP failures
- Remove mise.toml and internal planning doc from repo
- Update module docstrings to describe current purpose, not migration history
- Update tests to import from canonical relevance module
This commit is contained in:
Jeffrey Sperling
2026-03-11 18:40:07 -07:00
parent 6c402f66b7
commit 036bcd2ae3
18 changed files with 44 additions and 207 deletions
+1 -1
View File
@@ -58,7 +58,7 @@ def _extract_core_subject(topic: str) -> str:
Aggressively strip question/meta/research words to keep only the
core product/concept name (max 5 words).
"""
from .query import NOISE_WORDS, extract_core_subject
from .query import extract_core_subject
return extract_core_subject(topic, max_words=5, strip_suffixes=True)