fix(reddit): restore free path via keyless RSS + shreddit scrape (.json is dead) (#457)

* test(reddit): add live RSS + shreddit comment fixtures

Captured from reddit.com on 2026-05-29 (search.rss listing + the
/svc/shreddit/comments partial), trimmed to a representative subset plus
two synthetic edge cases (deleted author, negative score) for offline
parser tests.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(http): add keyless get_text helper

Browser-UA text fetch for RSS/HTML endpoints; returns None on any HTTP or
network failure so tiered callers fall through cleanly.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(reddit): keyless RSS discovery (search.rss + listing feeds)

Replaces the now-403 search.json with keyless Atom feeds, normalized to the
existing reddit_public post shape. Scores are placeholder zeros, backfilled
during shreddit enrichment.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(reddit): keyless shreddit comment scraper

Parses <shreddit-comment> elements from /svc/shreddit/comments/r/{sub}/t3_{id}
(score/author/created/permalink + thingId-anchored body) into top comments,
matching reddit_enrich output. Replaces the dead {thread}.json enrichment.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(reddit): tiered keyless orchestrator

Tier 0 one-shot .json (residential bonus) -> Tier 1 RSS discovery ->
Tier 2 shreddit enrichment. Returns [] never raises, so the SC backup
still engages when every keyless tier is empty.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reddit): route free path through keyless pipeline (.json is dead)

search_reddit_public is now a thin shim over reddit_keyless, so pipeline.py
and other callers need no change. Removes the dead .json enrichment helpers;
search/_parse_posts remain as the demoted Tier 0 attempt.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reddit): request sort=top so true top comments land on page 1

Guarantees the highest-scored comments are captured even on large threads,
independent of Reddit's default comment sort. Local score re-sort remains.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(reddit): recover post upvote scores via keyless listing partials

The shreddit community-more-posts partial server-renders each post's score
and comment count (works for normal users, not IP-gated), unlike RSS or the
comments endpoint. Use it as a scored discovery source and to backfill scores
onto RSS-discovered posts (subreddits derived from results when not provided).
Ranking now uses real upvote score.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reddit): listings backfill scores only on bare queries, not discovery

Caught running the full pipeline on a bare topic: deriving subreddits from
noisy RSS results and merging their top/hot listings flooded results with
high-upvote off-topic posts. Now derived-subreddit listings are used only to
backfill scores onto keyword-matched RSS posts; listing cards are merged as
discovery only when the caller explicitly provides subreddits (on-topic).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Matt Van Horn <455140+mvanhorn@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
Matt Van Horn
2026-05-29 14:43:56 -05:00
committed by GitHub
parent 1e03af19e0
commit 8d3a9e4368
14 changed files with 1389 additions and 216 deletions
+85
View File
@@ -0,0 +1,85 @@
"""Tests for scripts/lib/reddit_listing.py — keyless scored listing scrape."""
from pathlib import Path
from unittest import mock
from lib import reddit_listing as rl
FIXTURE = Path(__file__).resolve().parent.parent / "fixtures" / "reddit_listing_cards_sample.html"
def _html():
return FIXTURE.read_text(encoding="utf-8")
class TestParseCards:
"""parse_cards reads <shreddit-post> cards into scored post dicts."""
def test_parses_five_cards(self):
posts = rl.parse_cards(_html(), query="netherlands")
assert len(posts) == 5
def test_real_score_and_count(self):
posts = rl.parse_cards(_html())
top = posts[0]
assert top["score"] == 52692 # the real upvote count
assert top["engagement"]["score"] == 52692
assert top["num_comments"] == 1743
assert top["engagement"]["num_comments"] == 1743
def test_normalized_shape(self):
post = rl.parse_cards(_html())[0]
required = {"id", "title", "url", "score", "num_comments", "subreddit",
"created_utc", "author", "selftext", "date",
"engagement", "relevance", "why_relevant", "metadata"}
assert required.issubset(set(post.keys()))
assert post["why_relevant"] == "Reddit listing"
assert post["metadata"]["post_id"] # post id captured for backfill
def test_fields_populated(self):
post = rl.parse_cards(_html())[0]
assert post["title"]
assert post["author"] == "AdSpecialist6598"
assert post["subreddit"] == "technology"
assert "/comments/" in post["url"]
assert post["date"] and len(post["date"]) == 10
def test_empty_html_returns_empty(self):
assert rl.parse_cards("") == []
assert rl.parse_cards("<div>no cards</div>") == []
class TestListingUrl:
def test_top_includes_timeframe(self):
u = rl._listing_url("technology", "top")
assert "community-more-posts/top/" in u and "name=technology" in u and "t=month" in u
def test_hot_no_timeframe(self):
u = rl._listing_url("r/technology", "hot")
assert "community-more-posts/hot/" in u and "name=technology" in u and "t=" not in u
assert ".json" not in u
class TestFetchListings:
def test_dedupes_across_sorts(self):
with mock.patch.object(rl.http, "get_text", return_value=_html()):
posts = rl.fetch_listings(["technology"], depth="default")
urls = [p["url"] for p in posts]
assert len(urls) == len(set(urls)) # top + hot return same cards -> deduped
def test_no_subreddits_returns_empty(self):
assert rl.fetch_listings([], depth="default") == []
def test_all_fetches_fail_returns_empty(self):
with mock.patch.object(rl.http, "get_text", return_value=None):
assert rl.fetch_listings(["technology"]) == []
class TestScoreIndex:
def test_builds_post_id_to_score_map(self):
with mock.patch.object(rl.http, "get_text", return_value=_html()):
idx = rl.score_index(["technology"], depth="quick")
assert idx # non-empty
first = next(iter(idx.values()))
assert set(first.keys()) == {"score", "num_comments"}
assert any(v["score"] == 52692 for v in idx.values())