8d3a9e4368
* test(reddit): add live RSS + shreddit comment fixtures Captured from reddit.com on 2026-05-29 (search.rss listing + the /svc/shreddit/comments partial), trimmed to a representative subset plus two synthetic edge cases (deleted author, negative score) for offline parser tests. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(http): add keyless get_text helper Browser-UA text fetch for RSS/HTML endpoints; returns None on any HTTP or network failure so tiered callers fall through cleanly. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(reddit): keyless RSS discovery (search.rss + listing feeds) Replaces the now-403 search.json with keyless Atom feeds, normalized to the existing reddit_public post shape. Scores are placeholder zeros, backfilled during shreddit enrichment. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(reddit): keyless shreddit comment scraper Parses <shreddit-comment> elements from /svc/shreddit/comments/r/{sub}/t3_{id} (score/author/created/permalink + thingId-anchored body) into top comments, matching reddit_enrich output. Replaces the dead {thread}.json enrichment. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(reddit): tiered keyless orchestrator Tier 0 one-shot .json (residential bonus) -> Tier 1 RSS discovery -> Tier 2 shreddit enrichment. Returns [] never raises, so the SC backup still engages when every keyless tier is empty. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reddit): route free path through keyless pipeline (.json is dead) search_reddit_public is now a thin shim over reddit_keyless, so pipeline.py and other callers need no change. Removes the dead .json enrichment helpers; search/_parse_posts remain as the demoted Tier 0 attempt. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reddit): request sort=top so true top comments land on page 1 Guarantees the highest-scored comments are captured even on large threads, independent of Reddit's default comment sort. Local score re-sort remains. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(reddit): recover post upvote scores via keyless listing partials The shreddit community-more-posts partial server-renders each post's score and comment count (works for normal users, not IP-gated), unlike RSS or the comments endpoint. Use it as a scored discovery source and to backfill scores onto RSS-discovered posts (subreddits derived from results when not provided). Ranking now uses real upvote score. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reddit): listings backfill scores only on bare queries, not discovery Caught running the full pipeline on a bare topic: deriving subreddits from noisy RSS results and merging their top/hot listings flooded results with high-upvote off-topic posts. Now derived-subreddit listings are used only to backfill scores onto keyword-matched RSS posts; listing cards are merged as discovery only when the caller explicitly provides subreddits (on-topic). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Matt Van Horn <455140+mvanhorn@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
273 lines
9.2 KiB
Python
273 lines
9.2 KiB
Python
"""Reddit public ``.json`` search module (demoted to keyless Tier 0).
|
|
|
|
Reddit's public ``.json`` endpoints now return HTTP 403 from most contexts
|
|
(shreddit anti-bot), so this is no longer the primary free path. The keyless
|
|
pipeline (see reddit_keyless.py) still calls ``search`` as a cheap one-shot
|
|
Tier 0 attempt — a residential machine may occasionally get a 200 — before
|
|
falling through to RSS discovery (reddit_rss.py) and shreddit comment
|
|
enrichment (reddit_shreddit.py).
|
|
|
|
``search_reddit_public`` is retained as a compatibility shim that delegates to
|
|
the keyless pipeline, so existing callers (pipeline.py) need no change.
|
|
|
|
Endpoints (Tier 0):
|
|
- Global: https://www.reddit.com/search.json?q={query}&sort=relevance&t=month&limit={limit}
|
|
- Subreddit: https://www.reddit.com/r/{sub}/search.json?q={query}&restrict_sr=on&sort=relevance&t=month
|
|
|
|
Handles 429 rate limits with exponential backoff, HTML anti-bot responses,
|
|
network timeouts, and missing subreddits.
|
|
"""
|
|
|
|
import gzip
|
|
import json
|
|
import sys
|
|
import time
|
|
import urllib.error
|
|
import urllib.parse
|
|
import urllib.request
|
|
from typing import Any, Dict, List, Optional
|
|
|
|
|
|
USER_AGENT = (
|
|
"Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) "
|
|
"AppleWebKit/537.36 (KHTML, like Gecko) "
|
|
"Chrome/124.0.0.0 Safari/537.36"
|
|
)
|
|
|
|
# Depth-aware limits for thread counts
|
|
DEPTH_LIMITS = {
|
|
"quick": 10,
|
|
"default": 25,
|
|
"deep": 50,
|
|
}
|
|
|
|
MAX_RETRIES = 3
|
|
BASE_BACKOFF = 2.0 # seconds
|
|
|
|
|
|
def _log(msg: str):
|
|
"""Log to stderr."""
|
|
sys.stderr.write(f"[RedditPublic] {msg}\n")
|
|
sys.stderr.flush()
|
|
|
|
|
|
def _url_encode(text: str) -> str:
|
|
"""URL-encode a query string."""
|
|
return urllib.parse.quote_plus(text)
|
|
|
|
|
|
def _fetch_json(url: str, timeout: int = 15) -> Optional[Dict[str, Any]]:
|
|
"""Fetch JSON from a URL with retry on 429 and error handling.
|
|
|
|
Returns parsed JSON dict, or None on unrecoverable failure.
|
|
"""
|
|
headers = {
|
|
"User-Agent": USER_AGENT,
|
|
"Accept": "application/json",
|
|
"Accept-Language": "en-US,en;q=0.9",
|
|
"Accept-Encoding": "gzip, deflate",
|
|
"Connection": "keep-alive",
|
|
}
|
|
req = urllib.request.Request(url, headers=headers)
|
|
|
|
for attempt in range(MAX_RETRIES):
|
|
try:
|
|
with urllib.request.urlopen(req, timeout=timeout) as resp:
|
|
content_type = resp.headers.get("Content-Type", "")
|
|
if "json" not in content_type and "text/html" in content_type:
|
|
_log(f"Anti-bot HTML response (Content-Type: {content_type})")
|
|
return None
|
|
|
|
raw = resp.read()
|
|
if resp.headers.get("Content-Encoding", "").lower() == "gzip":
|
|
raw = gzip.decompress(raw)
|
|
body = raw.decode("utf-8")
|
|
return json.loads(body)
|
|
|
|
except urllib.error.HTTPError as e:
|
|
if e.code == 429:
|
|
delay = BASE_BACKOFF * (2 ** attempt)
|
|
retry_after = None
|
|
if hasattr(e, "headers"):
|
|
retry_after = e.headers.get("Retry-After")
|
|
if retry_after:
|
|
try:
|
|
delay = float(retry_after)
|
|
except ValueError:
|
|
pass
|
|
_log(f"429 rate limited, retry {attempt + 1}/{MAX_RETRIES} after {delay:.1f}s")
|
|
if attempt < MAX_RETRIES - 1:
|
|
time.sleep(delay)
|
|
continue
|
|
# Last attempt exhausted
|
|
_log("429 retries exhausted")
|
|
return None
|
|
elif e.code == 404:
|
|
_log(f"404 not found: {url}")
|
|
return None
|
|
elif e.code == 403:
|
|
_log(f"403 forbidden: {url}")
|
|
return None
|
|
else:
|
|
_log(f"HTTP {e.code}: {e.reason}")
|
|
return None
|
|
|
|
except (urllib.error.URLError, OSError, TimeoutError) as e:
|
|
_log(f"Network error: {e}")
|
|
return None
|
|
|
|
except json.JSONDecodeError as e:
|
|
_log(f"JSON decode error: {e}")
|
|
return None
|
|
|
|
return None
|
|
|
|
|
|
def _parse_posts(data: Optional[Dict[str, Any]]) -> List[Dict[str, Any]]:
|
|
"""Parse Reddit listing JSON into normalized post dicts."""
|
|
if not data:
|
|
return []
|
|
|
|
children = data.get("data", {}).get("children", [])
|
|
posts = []
|
|
|
|
for child in children:
|
|
if child.get("kind") != "t3":
|
|
continue
|
|
post = child.get("data", {})
|
|
permalink = str(post.get("permalink", "")).strip()
|
|
if not permalink or "/comments/" not in permalink:
|
|
continue
|
|
|
|
score = int(post.get("score", 0) or 0)
|
|
num_comments = int(post.get("num_comments", 0) or 0)
|
|
selftext = str(post.get("selftext", ""))
|
|
author = str(post.get("author", "[deleted]"))
|
|
created_utc = post.get("created_utc")
|
|
|
|
# Parse date
|
|
date_str = None
|
|
if created_utc:
|
|
try:
|
|
from datetime import datetime, timezone
|
|
dt = datetime.fromtimestamp(float(created_utc), tz=timezone.utc)
|
|
date_str = dt.strftime("%Y-%m-%d")
|
|
except (ValueError, TypeError, OSError):
|
|
pass
|
|
|
|
posts.append({
|
|
"id": "", # Will be assigned after dedup
|
|
"title": str(post.get("title", "")).strip(),
|
|
"url": f"https://www.reddit.com{permalink}",
|
|
"score": score,
|
|
"num_comments": num_comments,
|
|
"subreddit": str(post.get("subreddit", "")).strip(),
|
|
"created_utc": float(created_utc) if created_utc else None,
|
|
"author": author if author not in ("[deleted]", "[removed]") else "[deleted]",
|
|
"selftext": selftext[:500] if selftext else "",
|
|
# Normalized fields matching ScrapeCreators output
|
|
"date": date_str,
|
|
"engagement": {
|
|
"score": score,
|
|
"num_comments": num_comments,
|
|
"upvote_ratio": post.get("upvote_ratio"),
|
|
},
|
|
"relevance": _compute_relevance(score, num_comments),
|
|
"why_relevant": "Reddit public search",
|
|
"metadata": {},
|
|
})
|
|
|
|
return posts
|
|
|
|
|
|
def _compute_relevance(score: int, num_comments: int) -> float:
|
|
"""Estimate relevance from engagement signals."""
|
|
score_component = min(1.0, max(0.0, score / 500.0))
|
|
comments_component = min(1.0, max(0.0, num_comments / 200.0))
|
|
return round((score_component * 0.6) + (comments_component * 0.4), 3)
|
|
|
|
|
|
def search(
|
|
query: str,
|
|
depth: str = "default",
|
|
subreddit: Optional[str] = None,
|
|
timeout: int = 15,
|
|
) -> List[Dict[str, Any]]:
|
|
"""Search Reddit via the public JSON endpoint.
|
|
|
|
Args:
|
|
query: Search query string
|
|
depth: 'quick', 'default', or 'deep' — controls result limit
|
|
subreddit: Optional subreddit name (without r/) for scoped search
|
|
timeout: HTTP timeout in seconds
|
|
|
|
Returns:
|
|
List of normalized post dicts. Empty list on any failure.
|
|
"""
|
|
limit = DEPTH_LIMITS.get(depth, DEPTH_LIMITS["default"])
|
|
encoded_query = _url_encode(query)
|
|
|
|
if subreddit:
|
|
sub = subreddit.removeprefix("r/").strip()
|
|
url = (
|
|
f"https://www.reddit.com/r/{sub}/search.json"
|
|
f"?q={encoded_query}&restrict_sr=on&sort=relevance&t=month&limit={limit}&raw_json=1"
|
|
)
|
|
else:
|
|
url = (
|
|
f"https://www.reddit.com/search.json"
|
|
f"?q={encoded_query}&sort=relevance&t=month&limit={limit}&raw_json=1"
|
|
)
|
|
|
|
data = _fetch_json(url, timeout=timeout)
|
|
posts = _parse_posts(data)
|
|
|
|
# Dedupe by URL and assign IDs
|
|
seen_urls = set()
|
|
unique = []
|
|
for post in posts:
|
|
if post["url"] not in seen_urls:
|
|
seen_urls.add(post["url"])
|
|
unique.append(post)
|
|
|
|
for i, post in enumerate(unique):
|
|
post["id"] = f"R{i + 1}"
|
|
|
|
return unique[:limit]
|
|
|
|
|
|
def search_reddit_public(
|
|
topic: str,
|
|
from_date: str,
|
|
to_date: str,
|
|
depth: str = "default",
|
|
subreddits: Optional[List[str]] = None,
|
|
) -> List[Dict[str, Any]]:
|
|
"""High-level free Reddit search + enrichment (keyless).
|
|
|
|
Thin compatibility shim over the tiered keyless pipeline: the legacy
|
|
``.json`` search/enrichment endpoints now return HTTP 403, so this delegates
|
|
to ``reddit_keyless.search_and_enrich`` (Tier 0 one-shot ``.json`` →
|
|
Tier 1 RSS discovery → Tier 2 shreddit comment enrichment). The name and
|
|
signature are preserved so ``pipeline.py`` and other callers need no change
|
|
and the ScrapeCreators backup still engages when this returns empty.
|
|
|
|
The module-level ``search`` / ``_parse_posts`` helpers remain in use as the
|
|
keyless pipeline's demoted Tier 0 ``.json`` attempt.
|
|
|
|
Args:
|
|
topic: Search topic
|
|
from_date: Start date (YYYY-MM-DD)
|
|
to_date: End date (YYYY-MM-DD)
|
|
depth: 'quick', 'default', or 'deep'
|
|
subreddits: Optional list of subreddit names (without r/) for targeted search
|
|
|
|
Returns:
|
|
List of normalized item dicts matching ScrapeCreators output format.
|
|
Empty list on total failure (so SC backup can engage).
|
|
"""
|
|
from . import reddit_keyless
|
|
return reddit_keyless.search_and_enrich(
|
|
topic, from_date, to_date, depth=depth, subreddits=subreddits
|
|
)
|