4 Commits

Author SHA1 Message Date
Pnant 1ee30ecf28 fix: restore full cli file for xhs cookie import
ci / test (3.10) (push) Has been cancelled
ci / test (3.11) (push) Has been cancelled
ci / test (3.12) (push) Has been cancelled
ci / test (3.13) (push) Has been cancelled
2026-05-18 20:39:22 +08:00
Pnant 61d6536754 fix: support manual xhs cookies for xhs-cli (tests/test_cli.py) 2026-05-18 20:38:26 +08:00
Pnant 27b2cd2d02 fix: support manual xhs cookies for xhs-cli (agent_reach/guides/setup-xiaohongshu.md) 2026-05-18 20:38:24 +08:00
Pnant 7eae32d0b7 fix: support manual xhs cookies for xhs-cli (agent_reach/cli.py) 2026-05-18 20:38:21 +08:00
60 changed files with 2750 additions and 2695 deletions
+7
View File
@@ -0,0 +1,7 @@
{
"permissions": {
"allow": [
"WebFetch(domain:community.groq.com)"
]
}
}
-3
View File
@@ -13,6 +13,3 @@
# Groq Whisper (optional, for video transcription) — https://console.groq.com # Groq Whisper (optional, for video transcription) — https://console.groq.com
# GROQ_API_KEY=gsk_your_key_here # GROQ_API_KEY=gsk_your_key_here
# OpenAI Whisper (optional fallback when Groq is rate-limited) — https://platform.openai.com
# OPENAI_API_KEY=sk-your_key_here
-40
View File
@@ -28,43 +28,3 @@ jobs:
- name: Run tests - name: Run tests
run: | run: |
pytest -q pytest -q
# Editable installs (-e) never exercise wheel packaging, so a broken wheel
# can pass tests and still fail every real `pip install` from source.
# This job builds the actual wheel and installs it into a clean venv.
wheel-gate:
runs-on: ubuntu-latest
steps:
- name: Checkout
uses: actions/checkout@v4
- name: Setup Python
uses: actions/setup-python@v5
with:
python-version: "3.12"
- name: Build wheel
run: |
python -m pip install --upgrade pip build
python -m build
- name: Verify wheel has no duplicate entries and ships data files
run: |
python - <<'PY'
import glob, zipfile, collections
whl = glob.glob("dist/*.whl")[0]
names = zipfile.ZipFile(whl).namelist()
dupes = [n for n, c in collections.Counter(names).items() if c > 1]
assert not dupes, f"duplicate entries in wheel: {dupes}"
assert "agent_reach/skill/SKILL.md" in names, "SKILL.md missing from wheel"
for prefix in ("agent_reach/guides/", "agent_reach/scripts/", "agent_reach/skill/references/"):
assert any(n.startswith(prefix) for n in names), f"{prefix} missing from wheel"
print(f"wheel OK: {len(names)} entries, no duplicates, data files present")
PY
- name: Smoke-install wheel into clean venv
run: |
python -m venv /tmp/smoke
/tmp/smoke/bin/pip install --quiet dist/*.whl
/tmp/smoke/bin/agent-reach version
cd /tmp && /tmp/smoke/bin/python -c "import agent_reach; from importlib.resources import files; assert (files('agent_reach')/'skill'/'SKILL.md').is_file(); print('SKILL.md ships in site-packages OK')"
-4
View File
@@ -8,7 +8,3 @@ build/
.env .env
.agent-reach/ .agent-reach/
*.log *.log
# Claude Code personal permission settings — local only, never commit
.claude/settings.local.json
uv.lock
+2 -2
View File
@@ -1,9 +1,9 @@
# CLAUDE.md # CLAUDE.md
## Project ## Project
Agent Reach — Python CLI + library that gives AI agents read/search access to 13 internet platforms. Agent Reach — Python CLI + library that gives AI agents read/search access to 14+ internet platforms.
Positioning: installer + doctor + config tool. NOT a wrapper — after install, agents call upstream tools directly. Positioning: installer + doctor + config tool. NOT a wrapper — after install, agents call upstream tools directly.
Repo: github.com/Panniantong/Agent-Reach | License: MIT | Version: 1.4.2 Repo: github.com/Panniantong/Agent-Reach | License: MIT | Version: 1.3.0
## Commands ## Commands
- `pip install -e .` — Dev install - `pip install -e .` — Dev install
+46 -3
View File
@@ -6,7 +6,7 @@
<p align="center"> <p align="center">
<a href="LICENSE"><img src="https://img.shields.io/badge/License-MIT-blue.svg?style=for-the-badge" alt="MIT License"></a> <a href="LICENSE"><img src="https://img.shields.io/badge/License-MIT-blue.svg?style=for-the-badge" alt="MIT License"></a>
<a href="https://www.python.org/"><img src="https://img.shields.io/badge/Python-3.10+-green.svg?style=for-the-badge&logo=python&logoColor=white" alt="Python 3.10+"></a> <a href="https://www.python.org/"><img src="https://img.shields.io/badge/Python-3.10+-green.svg?style=for-the-badge&logo=python&logoColor=white" alt="Python 3.8+"></a>
<a href="https://github.com/Panniantong/agent-reach/stargazers"><img src="https://img.shields.io/github/stars/Panniantong/agent-reach?style=for-the-badge" alt="GitHub Stars"></a> <a href="https://github.com/Panniantong/agent-reach/stargazers"><img src="https://img.shields.io/github/stars/Panniantong/agent-reach?style=for-the-badge" alt="GitHub Stars"></a>
</p> </p>
@@ -75,7 +75,10 @@ AI Agent 已经能帮你写代码、改文档、管项目——但你让它去
| 📺 **B站** | 本地:字幕提取 + 搜索 | 服务器也能用 | 告诉 Agent「帮我配代理」 | | 📺 **B站** | 本地:字幕提取 + 搜索 | 服务器也能用 | 告诉 Agent「帮我配代理」 |
| 📖 **Reddit** | 搜索 + 读帖子和评论(通过 rdt-cli) | Cookie | 需要登录认证(`rdt login`),详见 [rdt-cli](https://github.com/public-clis/rdt-cli) | | 📖 **Reddit** | 搜索 + 读帖子和评论(通过 rdt-cli) | Cookie | 需要登录认证(`rdt login`),详见 [rdt-cli](https://github.com/public-clis/rdt-cli) |
| 📕 **小红书** | — | 阅读、搜索、发帖、评论、点赞 | 告诉 Agent「帮我配小红书」 | | 📕 **小红书** | — | 阅读、搜索、发帖、评论、点赞 | 告诉 Agent「帮我配小红书」 |
| 🎵 **抖音** | — | 视频解析、无水印下载链接获取 | 告诉 Agent「帮我配抖音」 |
| 💼 **LinkedIn** | Jina Reader 读公开页面 | Profile 详情、公司页面、职位搜索 | 告诉 Agent「帮我配 LinkedIn」 | | 💼 **LinkedIn** | Jina Reader 读公开页面 | Profile 详情、公司页面、职位搜索 | 告诉 Agent「帮我配 LinkedIn」 |
| 💬 **微信公众号** | 搜索 + 阅读公众号文章(全文 Markdown) | — | 无需配置 |
| 📰 **微博** | 热搜、搜索内容/用户/话题、用户动态、评论 | — | 无需配置 |
| 💻 **V2EX** | 热门帖子、节点帖子、帖子详情+回复、用户信息 | — | 无需配置 | | 💻 **V2EX** | 热门帖子、节点帖子、帖子详情+回复、用户信息 | — | 无需配置 |
| 📈 **雪球** | 股票行情、搜索股票、热门帖子、热门股票排行 | — | 告诉 Agent「帮我配雪球」 | | 📈 **雪球** | 股票行情、搜索股票、热门帖子、热门股票排行 | — | 告诉 Agent「帮我配雪球」 |
| 🎙️ **小宇宙播客** | — | 播客音频转文字(Whisper 转录,免费 Key) | 告诉 Agent「帮我配小宇宙播客」 | | 🎙️ **小宇宙播客** | — | 播客音频转文字(Whisper 转录,免费 Key) | 告诉 Agent「帮我配小宇宙播客」 |
@@ -172,7 +175,9 @@ channels/
├── bilibili.py → yt-dlp ← 可以换成 bilibili-api…… ├── bilibili.py → yt-dlp ← 可以换成 bilibili-api……
├── reddit.py → rdt-cli ← 搜索+阅读,需 Cookie 认证 ├── reddit.py → rdt-cli ← 搜索+阅读,需 Cookie 认证
├── xiaohongshu.py → mcporter MCP ← 可以换成其他 XHS 工具…… ├── xiaohongshu.py → mcporter MCP ← 可以换成其他 XHS 工具……
├── douyin.py → mcporter MCP ← 可以换成其他抖音工具……
├── linkedin.py → linkedin-mcp ← 可以换成 LinkedIn API…… ├── linkedin.py → linkedin-mcp ← 可以换成 LinkedIn API……
├── wechat.py → Exa (+ Camoufox) ← 搜索+阅读微信公众号文章
├── rss.py → feedparser ← 可以换成 atoma…… ├── rss.py → feedparser ← 可以换成 atoma……
├── exa_search.py → mcporter MCP ← 可以换成 Tavily、SerpAPI…… ├── exa_search.py → mcporter MCP ← 可以换成 Tavily、SerpAPI……
└── __init__.py → 渠道注册(doctor 检测用) └── __init__.py → 渠道注册(doctor 检测用)
@@ -193,10 +198,38 @@ channels/
| GitHub | [gh CLI](https://cli.github.com) | 官方工具,认证后完整 API 能力 | | GitHub | [gh CLI](https://cli.github.com) | 官方工具,认证后完整 API 能力 |
| 读 RSS | [feedparser](https://github.com/kurtmckee/feedparser) | Python 生态标准选择,2.3K Star | | 读 RSS | [feedparser](https://github.com/kurtmckee/feedparser) | Python 生态标准选择,2.3K Star |
| 小红书 | [xhs-cli](https://github.com/jackwener/xiaohongshu-cli) | 1.5K Starpipx 一行安装,搜索/阅读/评论/发帖 | | 小红书 | [xhs-cli](https://github.com/jackwener/xiaohongshu-cli) | 1.5K Starpipx 一行安装,搜索/阅读/评论/发帖 |
| 抖音 | [douyin-mcp-server](https://github.com/yzfly/douyin-mcp-server) | MCP 服务,无需登录,视频解析 + 无水印下载 |
| LinkedIn | [linkedin-scraper-mcp](https://github.com/stickerdaniel/linkedin-mcp-server) | ⭐1.2K,MCP 服务,浏览器自动化 | | LinkedIn | [linkedin-scraper-mcp](https://github.com/stickerdaniel/linkedin-mcp-server) | ⭐1.2K,MCP 服务,浏览器自动化 |
| 微信公众号 | [Exa](https://exa.ai)(搜索+阅读)+ [Camoufox](https://github.com/daijro/camoufox)(可选) | 零配置搜索+全文阅读,Camoufox 可选增强 |
> 📌 这些都是「当前选型」。不满意?换掉对应文件就行。这正是脚手架的意义。 > 📌 这些都是「当前选型」。不满意?换掉对应文件就行。这正是脚手架的意义。
### 抖音 / 小红书脚本提取的可选实现
如果你不只是想“解析抖音视频信息”,还想统一处理:
- 抖音视频脚本提取
- 小红书视频笔记脚本提取
- 小红书图文笔记正文 + 图片文字提取
- 固定输出 `script.md``info.json`
可以把 `douyin` 这个 mcporter alias 指向另一个兼容实现:
- [social-post-extractor-mcp](https://github.com/JNHFlow21/social-post-extractor-mcp)
这个实现保留了旧工具名兼容性:
- `parse_douyin_video_info`
- `get_douyin_download_link`
- `extract_douyin_text`
同时新增统一工具:
- `parse_social_post_info`
- `extract_social_post_script`
所以从 Agent Reach 的视角看,它依然只是一个 `mcporter` 里的 `douyin` server,只是能力更完整。
--- ---
## 安全性 ## 安全性
@@ -290,7 +323,7 @@ Agent Reach uses twitter-cli with cookie auth — zero API fees. Install with `p
<details> <details>
<summary><strong>Reddit 返回 403 怎么办?</strong></summary> <summary><strong>Reddit 返回 403 怎么办?</strong></summary>
Agent Reach 使用 [rdt-cli](https://github.com/public-clis/rdt-cli) 访问 Reddit。Reddit 自 2024 年起要求认证,安装后需运行 `rdt login` 登录。安装:`pipx install 'git+https://github.com/public-clis/rdt-cli.git'`(PyPI 版本暂时落后,从 GitHub 装),然后 `rdt login`(自动从浏览器提取 Cookie)。之后 Agent 可以用 `rdt search "关键词"` 搜索、`rdt read POST_ID` 读帖子全文和评论。 Agent Reach 使用 [rdt-cli](https://github.com/public-clis/rdt-cli) 访问 Reddit。Reddit 自 2024 年起要求认证,安装后需运行 `rdt login` 登录。安装:`pipx install rdt-cli`,然后 `rdt login`(自动从浏览器提取 Cookie)。之后 Agent 可以用 `rdt search "关键词"` 搜索、`rdt read POST_ID` 读帖子全文和评论。
</details> </details>
<details> <details>
@@ -305,6 +338,12 @@ Agent Reach 使用 [rdt-cli](https://github.com/public-clis/rdt-cli) 访问 Redd
安装 `pipx install xiaohongshu-cli`,然后 `xhs login`(自动从浏览器提取 Cookie)。之后 Agent 就能用 `xhs search "关键词"` 搜索笔记、`xhs read NOTE_ID` 阅读详情、`xhs comments NOTE_ID` 查看评论了。不需要 Docker。 安装 `pipx install xiaohongshu-cli`,然后 `xhs login`(自动从浏览器提取 Cookie)。之后 Agent 就能用 `xhs search "关键词"` 搜索笔记、`xhs read NOTE_ID` 阅读详情、`xhs comments NOTE_ID` 查看评论了。不需要 Docker。
</details> </details>
<details>
<summary><strong>怎么让 AI Agent 解析抖音视频?</strong></summary>
安装 douyin-mcp-server 后,Agent 就能用 `mcporter call 'douyin.parse_douyin_video_info(share_link: "分享链接")'` 解析视频信息、获取无水印下载链接。不需要登录,把抖音分享链接发给 Agent 就行。详见 https://github.com/yzfly/douyin-mcp-server
</details>
<details> <details>
<summary><strong>Compatible with Claude Code / Cursor / OpenClaw / Windsurf?</strong></summary> <summary><strong>Compatible with Claude Code / Cursor / OpenClaw / Windsurf?</strong></summary>
@@ -323,7 +362,7 @@ Yes! Agent Reach is an installer + configuration tool — any AI coding agent th
## 致谢 ## 致谢
[twitter-cli](https://github.com/public-clis/twitter-cli) · [rdt-cli](https://github.com/public-clis/rdt-cli) · [xhs-cli](https://github.com/jackwener/xiaohongshu-cli) · [bili-cli](https://github.com/public-clis/bilibili-cli) · [yt-dlp](https://github.com/yt-dlp/yt-dlp) · [Jina Reader](https://github.com/jina-ai/reader) · [Exa](https://exa.ai) · [mcporter](https://github.com/nicobailon/mcporter) · [feedparser](https://github.com/kurtmckee/feedparser) · [linkedin-scraper-mcp](https://github.com/stickerdaniel/linkedin-mcp-server) [twitter-cli](https://github.com/public-clis/twitter-cli) · [rdt-cli](https://github.com/public-clis/rdt-cli) · [xhs-cli](https://github.com/jackwener/xiaohongshu-cli) · [bili-cli](https://github.com/public-clis/bilibili-cli) · [yt-dlp](https://github.com/yt-dlp/yt-dlp) · [Jina Reader](https://github.com/jina-ai/reader) · [Exa](https://exa.ai) · [mcporter](https://github.com/nicobailon/mcporter) · [feedparser](https://github.com/kurtmckee/feedparser) · [douyin-mcp-server](https://github.com/yzfly/douyin-mcp-server) · [linkedin-scraper-mcp](https://github.com/stickerdaniel/linkedin-mcp-server)
## 联系 ## 联系
@@ -344,6 +383,10 @@ Yes! Agent Reach is an installer + configuration tool — any AI coding agent th
## 友情链接 ## 友情链接
[FluxNode](https://fluxnode.org) — 低价 AI API 中转站,官方一折,可按量或按套餐付费。可用于 OpenClaw、Claude Code 等一切 Agent。
[OpenClaw for Enterprise](https://github.com/littleben/openclaw-for-enterprise) — 企业级 OpenClaw 多用户部署方案,飞书里直接用 AI,容器隔离,一条命令管理。
[腾讯云 OpenClaw](https://www.tencentcloud.com/act/pro/intl-openclaw?referral_code=G76Y819A&lang=zh&pg=) — 在腾讯云Lighthouse秒级部署OpenClaw全能助手,可通过对话丝滑接入Agent Reach,给你的OpenClaw一键装上互联网能力。 [腾讯云 OpenClaw](https://www.tencentcloud.com/act/pro/intl-openclaw?referral_code=G76Y819A&lang=zh&pg=) — 在腾讯云Lighthouse秒级部署OpenClaw全能助手,可通过对话丝滑接入Agent Reach,给你的OpenClaw一键装上互联网能力。
## Star History ## Star History
+1 -1
View File
@@ -1,7 +1,7 @@
# -*- coding: utf-8 -*- # -*- coding: utf-8 -*-
"""Agent Reach — Give your AI Agent eyes to see the entire internet.""" """Agent Reach — Give your AI Agent eyes to see the entire internet."""
__version__ = "1.4.2" __version__ = "1.4.0"
__author__ = "Neo Reid" __author__ = "Neo Reid"
from agent_reach.core import AgentReach from agent_reach.core import AgentReach
-16
View File
@@ -1,16 +0,0 @@
# -*- coding: utf-8 -*-
"""Cross-channel backends.
A backend here is an upstream runtime that serves MULTIPLE channels
(e.g. OpenCLI covers xiaohongshu/reddit/bilibili/twitter through one
browser session), as opposed to the per-platform tools probed inside
each channel file.
"""
from .opencli import ( # noqa: F401
OPENCLI_EXTENSION_URL,
OPENCLI_PACKAGE,
OpenCLIStatus,
opencli_status,
opencli_summary,
)
-136
View File
@@ -1,136 +0,0 @@
# -*- coding: utf-8 -*-
"""OpenCLI backend probing.
OpenCLI (github.com/jackwener/opencli) drives the user's real Chrome via a
browser-bridge extension + local daemon, reusing existing login sessions —
zero per-platform configuration, desktop-only (no headless).
Probing notes (verified live):
- `opencli doctor` AUTO-STARTS the daemon — a side effect, so health
checks must use `opencli daemon status` (pure query) instead.
- Exit codes are always 0; status must be parsed from text output.
- "Extension: disconnected" does NOT mean unusable: the extension's
service worker sleeps and any real opencli command wakes it up
(verified: status flips disconnected→connected after one call).
Since daemon status can't tell "sleeping" from "never installed",
we check Chrome's Extensions directory on disk to disambiguate.
"""
import glob
import os
from dataclasses import dataclass
from agent_reach.probe import probe_command
OPENCLI_PACKAGE = "@jackwener/opencli"
OPENCLI_EXTENSION_ID = "ildkmabpimmkaediidaifkhjpohdnifk"
OPENCLI_EXTENSION_URL = (
f"https://chromewebstore.google.com/detail/opencli/{OPENCLI_EXTENSION_ID}"
)
#: Chrome-family profile roots that contain <Profile>/Extensions/<id>/
_CHROME_PROFILE_ROOTS = (
"~/Library/Application Support/Google/Chrome", # macOS Chrome
"~/Library/Application Support/Chromium", # macOS Chromium
"~/.config/google-chrome", # Linux Chrome
"~/.config/chromium", # Linux Chromium
)
def _extension_installed_on_disk() -> bool:
"""True if the OpenCLI extension exists in any Chrome profile.
Store-installed extensions always live under
<profile>/Extensions/<extension id>/ — this disambiguates a sleeping
service worker from a never-installed extension. Dev installs via
"Load unpacked" are not covered (those users can read `opencli doctor`).
"""
roots = [os.path.expanduser(p) for p in _CHROME_PROFILE_ROOTS]
local_app_data = os.environ.get("LOCALAPPDATA")
if local_app_data: # Windows
roots.append(os.path.join(local_app_data, "Google", "Chrome", "User Data"))
for root in roots:
if glob.glob(os.path.join(root, "*", "Extensions", OPENCLI_EXTENSION_ID)):
return True
return False
@dataclass
class OpenCLIStatus:
installed: bool = False
broken: bool = False
daemon_running: bool = False
extension_connected: bool = False
extension_installed: bool = False
version: str = ""
hint: str = ""
@property
def ready(self) -> bool:
"""Usable now or on first call.
A live connection counts, and so does an installed-but-sleeping
extension: its service worker wakes on the first real command.
"""
return self.installed and not self.broken and (
self.extension_connected or self.extension_installed
)
def opencli_status(timeout: int = 10) -> OpenCLIStatus:
"""Probe OpenCLI install + daemon/extension state without side effects."""
version_probe = probe_command(
"opencli", ["--version"], timeout=timeout, package=OPENCLI_PACKAGE
)
if version_probe.status == "missing":
return OpenCLIStatus(installed=False)
if not version_probe.ok:
return OpenCLIStatus(
installed=True,
broken=True,
hint=(
"opencli 命令存在但无法执行(node 环境损坏),重装:\n"
f" npm install -g {OPENCLI_PACKAGE}"
),
)
st = OpenCLIStatus(installed=True, version=version_probe.output.strip())
daemon_probe = probe_command(
"opencli", ["daemon", "status"], timeout=timeout, package=OPENCLI_PACKAGE
)
output = daemon_probe.output if daemon_probe.ok else ""
# `opencli daemon status` prints lines like:
# Daemon: running (PID 37389) / Daemon: not running
# Extension: connected / Extension: disconnected
for line in output.splitlines():
line = line.strip().lower()
if line.startswith("daemon:"):
st.daemon_running = "not running" not in line and "running" in line
elif line.startswith("extension:"):
st.extension_connected = "disconnected" not in line and "connected" in line
if not st.extension_connected:
st.extension_installed = _extension_installed_on_disk()
if not st.extension_installed:
st.hint = (
"OpenCLI 已安装,但 Chrome 扩展未安装。\n"
f" 1. 安装扩展(需手动点一次):{OPENCLI_EXTENSION_URL}\n"
" 2. 保持 Chrome 打开,运行 `opencli doctor` 验证"
)
return st
def opencli_summary(st: OpenCLIStatus) -> str:
"""One-line state description for channel messages / install output."""
if not st.installed:
return "OpenCLI 未安装"
if st.broken:
return "OpenCLI 无法执行(node 环境损坏)"
if st.extension_connected:
return f"OpenCLI 可用(浏览器登录态,v{st.version}"
if st.ready:
return "OpenCLI 可用(扩展睡眠中,调用时自动唤醒)"
if st.daemon_running:
return "OpenCLI 已安装,等待 Chrome 扩展安装"
return "OpenCLI 已安装(daemon 未运行,使用时自动启动;需 Chrome 扩展)"
+6
View File
@@ -16,7 +16,10 @@ from .rss import RSSChannel
from .bilibili import BilibiliChannel from .bilibili import BilibiliChannel
from .exa_search import ExaSearchChannel from .exa_search import ExaSearchChannel
from .xiaohongshu import XiaoHongShuChannel from .xiaohongshu import XiaoHongShuChannel
from .douyin import DouyinChannel
from .linkedin import LinkedInChannel from .linkedin import LinkedInChannel
from .wechat import WeChatChannel
from .weibo import WeiboChannel
from .xiaoyuzhou import XiaoyuzhouChannel from .xiaoyuzhou import XiaoyuzhouChannel
from .v2ex import V2EXChannel from .v2ex import V2EXChannel
from .xueqiu import XueqiuChannel from .xueqiu import XueqiuChannel
@@ -29,7 +32,10 @@ ALL_CHANNELS: List[Channel] = [
RedditChannel(), RedditChannel(),
BilibiliChannel(), BilibiliChannel(),
XiaoHongShuChannel(), XiaoHongShuChannel(),
DouyinChannel(),
LinkedInChannel(), LinkedInChannel(),
WeChatChannel(),
WeiboChannel(),
XiaoyuzhouChannel(), XiaoyuzhouChannel(),
V2EXChannel(), V2EXChannel(),
XueqiuChannel(), XueqiuChannel(),
+3 -37
View File
@@ -8,22 +8,11 @@ and provides:
- check(config) → is the upstream tool installed and configured? - check(config) → is the upstream tool installed and configured?
After installation, agents call upstream tools directly. After installation, agents call upstream tools directly.
Backend routing semantics:
- `backends` is an ORDERED candidate list: backends[0] is the preferred
backend, the rest are fallbacks. "Switching backends" for a platform
means reordering this list (or a user override) — not rewriting code.
- check() must set `self.active_backend` to the backend that is actually
serving the channel right now (None when nothing usable is found).
shutil.which() alone is NOT proof of health — a stale venv shim passes
which() but cannot execute (see agent_reach.probe). Channels should
really execute a lightweight command before claiming a backend active.
- Users can force a backend with config key `<channel>_backend`
(or env var `<CHANNEL>_BACKEND`); ordered_backends() applies it.
""" """
import shutil
from abc import ABC, abstractmethod from abc import ABC, abstractmethod
from typing import List, Optional, Tuple from typing import List, Tuple
class Channel(ABC): class Channel(ABC):
@@ -31,40 +20,17 @@ class Channel(ABC):
name: str = "" # e.g. "youtube" name: str = "" # e.g. "youtube"
description: str = "" # e.g. "YouTube 视频和字幕" description: str = "" # e.g. "YouTube 视频和字幕"
backends: List[str] = [] # ordered candidates — backends[0] = preferred backends: List[str] = [] # e.g. ["yt-dlp"] — what upstream tool is used
tier: int = 0 # 0=zero-config, 1=needs free key, 2=needs setup tier: int = 0 # 0=zero-config, 1=needs free key, 2=needs setup
#: Backend currently serving this channel; set by check(), None = unavailable.
active_backend: Optional[str] = None
@abstractmethod @abstractmethod
def can_handle(self, url: str) -> bool: def can_handle(self, url: str) -> bool:
"""Check if this channel can handle this URL.""" """Check if this channel can handle this URL."""
... ...
def ordered_backends(self, config=None) -> List[str]:
"""Candidate backends in probe order, honoring the user override.
The config key `<channel>_backend` (env `<CHANNEL>_BACKEND`) moves the
named backend to the front of the list; unknown values are ignored so
a stale override can never hide working backends.
"""
candidates = list(self.backends)
override = config.get(f"{self.name}_backend") if config else None
if override:
for i, b in enumerate(candidates):
if b == override or b.startswith(override):
candidates.insert(0, candidates.pop(i))
break
return candidates
def check(self, config=None) -> Tuple[str, str]: def check(self, config=None) -> Tuple[str, str]:
""" """
Check if this channel's upstream tool is available. Check if this channel's upstream tool is available.
Returns (status, message) where status is 'ok'/'warn'/'off'/'error'. Returns (status, message) where status is 'ok'/'warn'/'off'/'error'.
Subclasses with external backends must really probe them (see
agent_reach.probe.probe_command) and set self.active_backend.
""" """
self.active_backend = self.backends[0] if self.backends else "内置"
return "ok", f"{''.join(self.backends) if self.backends else '内置'}" return "ok", f"{''.join(self.backends) if self.backends else '内置'}"
+8 -27
View File
@@ -3,10 +3,9 @@
import json import json
import os import os
import shutil
import subprocess
import urllib.request import urllib.request
from agent_reach.probe import probe_command
from .base import Channel from .base import Channel
_UA = "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36" _UA = "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36"
@@ -37,23 +36,11 @@ class BilibiliChannel(Channel):
return "bilibili.com" in d or "b23.tv" in d return "bilibili.com" in d or "b23.tv" in d
def check(self, config=None): def check(self, config=None):
self.active_backend = None if not shutil.which("yt-dlp"):
# 真跑 yt-dlp --version,区分 未装 / 断链 / 异常(which 命中不等于能用)
yt = probe_command("yt-dlp", ["--version"], timeout=10, package="yt-dlp")
if yt.status == "missing":
return "off", "yt-dlp 未安装。安装:pip install yt-dlp" return "off", "yt-dlp 未安装。安装:pip install yt-dlp"
if yt.status == "broken":
return "error", "yt-dlp 已安装但无法执行\n" + yt.hint
if not yt.ok:
detail = yt.hint or yt.output or yt.status
return "error", f"yt-dlp 探测失败({yt.status}):{detail}"
self.active_backend = "yt-dlp"
proxy = (config.get("bilibili_proxy") if config else None) or os.environ.get("BILIBILI_PROXY") proxy = (config.get("bilibili_proxy") if config else None) or os.environ.get("BILIBILI_PROXY")
# 真跑 bili --version——断链时旧的 which 检测会误报"bili-cli 可用" has_bili_cli = bool(shutil.which("bili"))
bili = probe_command("bili", ["--version"], timeout=10, package="bilibili-cli")
parts = [] parts = []
@@ -64,22 +51,16 @@ class BilibiliChannel(Channel):
parts.append("视频读取:yt-dlp") parts.append("视频读取:yt-dlp")
# bili-cli 增强 # bili-cli 增强
if bili.ok: if has_bili_cli:
parts.append("搜索/热门/排行:bili-cli 可用") parts.append("搜索/热门/排行:bili-cli 可用")
status = "ok"
else: else:
if bili.status == "broken": # 检测搜索 API 连通性
parts.append("bili-cli 已安装但无法执行,不计为可用\n" + bili.hint)
elif bili.status in ("timeout", "error"):
parts.append(f"bili-cli 探测失败({bili.status}),不计为可用")
# 降级走搜索 API;只探测一次,message 和 status 共用结果
api_ok = _search_api_ok() api_ok = _search_api_ok()
if api_ok: if api_ok:
parts.append("搜索:B站 API 可用") parts.append("搜索:B站 API 可用")
else: else:
parts.append("搜索:B站 API 不可达") parts.append("搜索:B站 API 不可达")
if bili.status == "missing": parts.append("提示:安装 bili-cli 可解锁热门/排行/动态:pipx install bilibili-cli")
parts.append("提示:安装 bili-cli 可解锁热门/排行/动态:pipx install bilibili-cli")
status = "ok" if api_ok else "warn"
status = "ok" if has_bili_cli or _search_api_ok() else "warn"
return status, "".join(parts) return status, "".join(parts)
+56
View File
@@ -0,0 +1,56 @@
# -*- coding: utf-8 -*-
"""Douyin (抖音) — check if mcporter + douyin-mcp-server is available."""
import shutil
import subprocess
from .base import Channel
class DouyinChannel(Channel):
name = "douyin"
description = "抖音短视频"
backends = ["douyin-mcp-server"]
tier = 2
def can_handle(self, url: str) -> bool:
from urllib.parse import urlparse
d = urlparse(url).netloc.lower()
return "douyin.com" in d or "iesdouyin.com" in d
def check(self, config=None):
mcporter = shutil.which("mcporter")
if not mcporter:
return "off", (
"需要 mcporter + douyin-mcp-server。安装步骤:\n"
" 1. npm install -g mcporter\n"
" 2. pip install douyin-mcp-server\n"
" 3. 启动服务(见下方说明)\n"
" 4. mcporter config add douyin http://localhost:18070/mcp\n"
" 详见 https://github.com/yzfly/douyin-mcp-server"
)
try:
r = subprocess.run(
[mcporter, "config", "list"], capture_output=True,
encoding="utf-8", errors="replace", timeout=5
)
if "douyin" not in r.stdout:
return "off", (
"mcporter 已装但抖音 MCP 未配置。运行:\n"
" pip install douyin-mcp-server\n"
" # 启动服务后:\n"
" mcporter config add douyin http://localhost:18070/mcp"
)
except Exception:
return "off", "mcporter 连接异常"
# Verify MCP connectivity by listing available tools instead of
# calling with a hardcoded (invalid) share link that always fails.
try:
r = subprocess.run(
[mcporter, "list", "douyin"],
capture_output=True, encoding="utf-8", errors="replace", timeout=15
)
if r.returncode == 0 and r.stdout.strip():
return "ok", "完整可用(视频解析、下载链接获取)"
return "warn", "MCP 已连接但工具列表为空,检查 douyin-mcp-server 服务是否在运行"
except Exception:
return "warn", "MCP 连接异常,检查 douyin-mcp-server 服务是否在运行"
+17 -19
View File
@@ -1,13 +1,10 @@
# -*- coding: utf-8 -*- # -*- coding: utf-8 -*-
"""Exa Search — check if mcporter + Exa MCP is available.""" """Exa Search — check if mcporter + Exa MCP is available."""
from agent_reach.probe import probe_command import shutil
import subprocess
from .base import Channel from .base import Channel
#: mcporter 是 npm 包,断链处方与默认的 pipx/uv 不同
_MCPORTER_BROKEN_HINT = "mcporter 无法执行(node 环境损坏),重装:\n npm install -g mcporter"
class ExaSearchChannel(Channel): class ExaSearchChannel(Channel):
name = "exa_search" name = "exa_search"
@@ -19,22 +16,23 @@ class ExaSearchChannel(Channel):
return False # Search-only channel return False # Search-only channel
def check(self, config=None): def check(self, config=None):
self.active_backend = None mcporter = shutil.which("mcporter")
probe = probe_command("mcporter", ["config", "list"], timeout=10, package="mcporter") if not mcporter:
if probe.status == "missing":
return "off", ( return "off", (
"需要 mcporter + Exa MCP。安装:\n" "需要 mcporter + Exa MCP。安装:\n"
" npm install -g mcporter\n" " npm install -g mcporter\n"
" mcporter config add exa https://mcp.exa.ai/mcp" " mcporter config add exa https://mcp.exa.ai/mcp"
) )
if probe.status == "broken": try:
return "error", _MCPORTER_BROKEN_HINT r = subprocess.run(
if not probe.ok: # timeout / error [mcporter, "config", "list"], capture_output=True,
return "error", f"mcporter 执行异常:{probe.hint or probe.output or probe.status}" encoding="utf-8", errors="replace", timeout=5
if "exa" in probe.output.lower(): )
self.active_backend = self.backends[0] if "exa" in r.stdout.lower():
return "ok", "全网语义搜索可用(免费,无需 API Key)" return "ok", "全网语义搜索可用(免费,无需 API Key)"
return "off", ( return "off", (
"mcporter 已装但 Exa 未配置。运行:\n" "mcporter 已装但 Exa 未配置。运行:\n"
" mcporter config add exa https://mcp.exa.ai/mcp" " mcporter config add exa https://mcp.exa.ai/mcp"
) )
except Exception:
return "off", "mcporter 连接异常"
+13 -23
View File
@@ -1,8 +1,8 @@
# -*- coding: utf-8 -*- # -*- coding: utf-8 -*-
"""GitHub — check if gh CLI is available.""" """GitHub — check if gh CLI is available."""
from agent_reach.probe import probe_command import shutil
import subprocess
from .base import Channel from .base import Channel
@@ -17,26 +17,16 @@ class GitHubChannel(Channel):
return "github.com" in urlparse(url).netloc.lower() return "github.com" in urlparse(url).netloc.lower()
def check(self, config=None): def check(self, config=None):
# 真跑 gh auth status 探活。注意:未登录时 rc!=0 是正常业务态(warn),不是 error。 gh = shutil.which("gh")
probe = probe_command("gh", ["auth", "status"], timeout=10, package="gh") if not gh:
if probe.status == "missing":
self.active_backend = None
return "warn", "gh CLI 未安装。安装:https://cli.github.com" return "warn", "gh CLI 未安装。安装:https://cli.github.com"
if probe.status == "broken": try:
# gh 是二进制安装(brew/官方包),不是 pip 包——处方不用 pipx/uv 文案 r = subprocess.run(
self.active_backend = None [gh, "auth", "status"],
return "error", ( capture_output=True, encoding="utf-8", errors="replace", timeout=5
"gh 命令存在但无法执行——安装已损坏。重装即可修复:\n"
" brew reinstall gh\n"
"或从 https://cli.github.com 重新安装 gh CLI"
) )
if probe.status == "timeout": if r.returncode == 0:
# gh 本体能启动(工具是活的),只是状态检查超时 return "ok", "完整可用(读取、搜索、Fork、Issue、PR 等)"
self.active_backend = "gh CLI" return "warn", "gh CLI 已安装但未认证。运行 gh auth login 可解锁完整功能"
return "warn", "gh CLI 状态检查超时,运行 gh auth status 查看详情" except Exception:
if probe.ok: return "warn", "gh CLI 状态检查失败,运行 gh auth status 查看详情"
self.active_backend = "gh CLI"
return "ok", "完整可用(读取、搜索、Fork、Issue、PR 等)"
# rc != 0gh 活着但未认证(gh auth status 的正常业务态)
self.active_backend = "gh CLI"
return "warn", "gh CLI 已安装但未认证。运行 gh auth login 可解锁完整功能"
+13 -15
View File
@@ -1,13 +1,10 @@
# -*- coding: utf-8 -*- # -*- coding: utf-8 -*-
"""LinkedIn — check if linkedin-scraper-mcp is available.""" """LinkedIn — check if linkedin-scraper-mcp is available."""
from agent_reach.probe import probe_command import shutil
import subprocess
from .base import Channel from .base import Channel
#: mcporter 是 npm 包,断链处方与默认的 pipx/uv 不同
_MCPORTER_BROKEN_HINT = "mcporter 无法执行(node 环境损坏),重装:\n npm install -g mcporter"
class LinkedInChannel(Channel): class LinkedInChannel(Channel):
name = "linkedin" name = "linkedin"
@@ -20,22 +17,23 @@ class LinkedInChannel(Channel):
return "linkedin.com" in urlparse(url).netloc.lower() return "linkedin.com" in urlparse(url).netloc.lower()
def check(self, config=None): def check(self, config=None):
self.active_backend = None mcporter = shutil.which("mcporter")
probe = probe_command("mcporter", ["config", "list"], timeout=10, package="mcporter") if not mcporter:
if probe.status == "missing":
return "off", ( return "off", (
"基本内容可通过 Jina Reader 读取。完整功能需要:\n" "基本内容可通过 Jina Reader 读取。完整功能需要:\n"
" pip install linkedin-scraper-mcp\n" " pip install linkedin-scraper-mcp\n"
" mcporter config add linkedin http://localhost:3000/mcp\n" " mcporter config add linkedin http://localhost:3000/mcp\n"
" 详见 https://github.com/stickerdaniel/linkedin-mcp-server" " 详见 https://github.com/stickerdaniel/linkedin-mcp-server"
) )
if probe.status == "broken": try:
return "error", _MCPORTER_BROKEN_HINT r = subprocess.run(
if not probe.ok: # timeout / error [mcporter, "config", "list"], capture_output=True,
return "error", f"mcporter 执行异常:{probe.hint or probe.output or probe.status}" encoding="utf-8", errors="replace", timeout=5
if "linkedin" in probe.output.lower(): )
self.active_backend = "linkedin-scraper-mcp" if "linkedin" in r.stdout.lower():
return "ok", "完整可用(Profile、公司、职位搜索)" return "ok", "完整可用(Profile、公司、职位搜索)"
except Exception:
pass
return "off", ( return "off", (
"mcporter 已装但 LinkedIn MCP 未配置。运行:\n" "mcporter 已装但 LinkedIn MCP 未配置。运行:\n"
" pip install linkedin-scraper-mcp\n" " pip install linkedin-scraper-mcp\n"
+28 -73
View File
@@ -10,30 +10,16 @@ import json
import shutil import shutil
import subprocess import subprocess
from agent_reach.utils.process import utf8_subprocess_env
from .base import Channel from .base import Channel
_CREDENTIAL_FILE = "~/.config/rdt-cli/credential.json" _CREDENTIAL_FILE = "~/.config/rdt-cli/credential.json"
# Pinned to the 0.4.2 state — PyPI still only has 0.4.1 (upstream issue #10).
_RDT_GIT_SOURCE = "git+https://github.com/public-clis/rdt-cli.git@5e4fb3720d5c174e976cd425ccc3b879d52cac66"
#: shell 对"找到但不可执行/找不到"使用的退出码(对齐 agent_reach.probe
_BROKEN_EXIT_CODES = (126, 127)
#: rdt 应从固定 git 源安装(PyPI 落后),断链处方与 probe 默认的 pipx/uv 不同
_RDT_BROKEN_HINT = (
"rdt 命令存在但无法执行——通常是系统 Python 升级后 venv 解释器丢失。\n"
"PyPI 版本落后,推荐用固定 git 源强制重装:\n"
f" pipx install --force '{_RDT_GIT_SOURCE}'"
)
class RedditChannel(Channel): class RedditChannel(Channel):
name = "reddit" name = "reddit"
description = "Reddit 帖子和评论" description = "Reddit 帖子和评论"
backends = ["rdt-cli"] backends = ["rdt-cli"]
tier = 1 # Reddit requires login since 2024 (rdt login) — not zero-config tier = 0
def can_handle(self, url: str) -> bool: def can_handle(self, url: str) -> bool:
from urllib.parse import urlparse from urllib.parse import urlparse
@@ -42,24 +28,17 @@ class RedditChannel(Channel):
return "reddit.com" in d or "redd.it" in d return "reddit.com" in d or "redd.it" in d
def check(self, config=None): def check(self, config=None):
self.active_backend = None
rdt = shutil.which("rdt") rdt = shutil.which("rdt")
if not rdt: if not rdt:
return "off", ( return "off", (
"需要安装 rdt-cli。PyPI 版本可能暂时落后,推荐直接从 GitHub 安装\n" "需要安装 rdt-cli(推荐使用最新版 v0.4.2+\n"
f" pipx install '{_RDT_GIT_SOURCE}'\n" " pip install 'rdt-cli>=0.4.2'\n"
"如已确认 PyPI 版本已更新,也可使用\n" "\n"
" pipx install rdt-cli\n"
" uv tool install rdt-cli\n" " uv tool install rdt-cli\n"
"最新源码:https://github.com/public-clis/rdt-cli\n" "最新源码:https://github.com/public-clis/rdt-cli\n"
"安装后运行 `rdt login` 登录(需先在浏览器登录 reddit.com)" "安装后运行 `rdt login` 登录(需先在浏览器登录 reddit.com)"
) )
# 不走 probe_command:实测 `rdt status --json` 成功时(rc=0)也会向 stderr
# 打网络重试日志,probe 把 stdout+stderr 合并后 JSON 解析必炸。
# 故保留手写 subprocess(stdout 单独捕获),但异常分类对齐 probe 语义:
# exec 失败/126/127 → brokenvenv 断链处方),TimeoutExpired → 超时。
try: try:
r = subprocess.run( r = subprocess.run(
[rdt, "status", "--json"], [rdt, "status", "--json"],
@@ -67,55 +46,31 @@ class RedditChannel(Channel):
encoding="utf-8", encoding="utf-8",
errors="replace", errors="replace",
timeout=10, timeout=10,
env=utf8_subprocess_env(),
) )
except subprocess.TimeoutExpired: data = json.loads(r.stdout or "{}")
return "error", "rdt 响应超时(>10s),Reddit 状态未知。稍后重试或运行 `rdt status` 查看详情" authenticated = data.get("data", {}).get("authenticated", False)
except OSError: username = data.get("data", {}).get("username") or ""
# 含 FileNotFoundErrorwhich 命中但 exec 失败 = venv 断链(probe 的 broken
return "error", _RDT_BROKEN_HINT
if r.returncode in _BROKEN_EXIT_CODES: if authenticated:
return "error", _RDT_BROKEN_HINT suffix = f"(已登录:{username}" if username else ""
return "ok", (f"rdt-cli 可用{suffix}(搜索帖子、阅读全文、查看评论)")
if r.returncode != 0: return "warn", (
detail = (r.stderr or r.stdout or "").strip().splitlines() "rdt-cli 已安装但未登录。Reddit 自 2024 年起要求认证,"
tail = detail[-1] if detail else "无输出" "未登录时所有请求均返回 403。\n\n"
return "error", f"rdt 异常退出(exit {r.returncode}):{tail}。运行 `rdt status` 查看详情" "方法一(自动):运行 `rdt login`\n"
" 先在浏览器登录 reddit.com,再运行此命令自动提取 Cookie。\n\n"
"方法二(手动,适用于 Chrome/Edge 127+ 无法自动提取时):\n"
" 1. Chrome 应用商店安装 Cookie-Editor 扩展:\n"
" https://chromewebstore.google.com/detail/cookie-editor/hlkenndednhfkekhgcdicdfddnkalmdm\n"
" 2. 在浏览器打开 reddit.com(确保已登录)\n"
" 3. 点击 Cookie-Editor 图标,找到 `reddit_session`,复制其 Value\n"
f" 4. 将以下内容写入 {_CREDENTIAL_FILE}\n"
' {"cookies": {"reddit_session": "<粘贴 Value>"}, '
'"source": "manual", "username": "<你的用户名>", '
'"modhash": null, "saved_at": 0, "last_verified_at": null}\n\n'
"验证:`rdt status --json` 确认 authenticated: true"
)
# 进程正常退出 → rdt 本身是活的(无论登录与否),后端即为可用 except (json.JSONDecodeError, FileNotFoundError, subprocess.TimeoutExpired):
self.active_backend = "rdt-cli" return "warn", "rdt-cli 已安装但状态检查失败,运行 `rdt status` 查看详情"
try:
data = json.loads(r.stdout or "")
except json.JSONDecodeError:
data = None
if not isinstance(data, dict):
return "warn", "rdt-cli 可用但状态输出无法解析,运行 `rdt status` 查看登录状态"
info = data.get("data")
if not isinstance(info, dict):
info = {}
authenticated = info.get("authenticated", False)
username = info.get("username") or ""
if authenticated:
suffix = f"(已登录:{username}" if username else ""
return "ok", (f"rdt-cli 可用{suffix}(搜索帖子、阅读全文、查看评论)")
return "warn", (
"rdt-cli 已安装但未登录。Reddit 自 2024 年起要求认证,"
"未登录时所有请求均返回 403。\n\n"
"方法一(自动):运行 `rdt login`\n"
" 先在浏览器登录 reddit.com,再运行此命令自动提取 Cookie。\n\n"
"方法二(手动,适用于 Chrome/Edge 127+ 无法自动提取时):\n"
" 1. Chrome 应用商店安装 Cookie-Editor 扩展:\n"
" https://chromewebstore.google.com/detail/cookie-editor/hlkenndednhfkekhgcdicdfddnkalmdm\n"
" 2. 在浏览器打开 reddit.com(确保已登录)\n"
" 3. 点击 Cookie-Editor 图标,找到 `reddit_session`,复制其 Value\n"
f" 4. 将以下内容写入 {_CREDENTIAL_FILE}\n"
' {"cookies": {"reddit_session": "<粘贴 Value>"}, '
'"source": "manual", "username": "<你的用户名>", '
'"modhash": null, "saved_at": 0, "last_verified_at": null}\n\n'
"验证:`rdt status --json` 确认 authenticated: true"
)
+2 -8
View File
@@ -15,13 +15,7 @@ class RSSChannel(Channel):
def check(self, config=None): def check(self, config=None):
try: try:
import feedparser # noqa: F401 import feedparser
return "ok", "可读取 RSS/Atom 源"
except ImportError: except ImportError:
self.active_backend = None
return "off", "feedparser 未安装。安装:pip install feedparser" return "off", "feedparser 未安装。安装:pip install feedparser"
except Exception as e:
# 已安装但导入期崩溃(半残安装/版本冲突)→ 重装处方
self.active_backend = None
return "error", f"feedparser 导入失败:{e}\n修复:pip install --force-reinstall feedparser"
self.active_backend = self.backends[0]
return "ok", "可读取 RSS/Atom 源"
+46 -86
View File
@@ -1,8 +1,9 @@
# -*- coding: utf-8 -*- # -*- coding: utf-8 -*-
"""Twitter/X — check if twitter-cli or bird CLI is available.""" """Twitter/X — check if twitter-cli or bird CLI is available."""
import shutil
import subprocess
from .base import Channel from .base import Channel
from agent_reach.probe import probe_command
class TwitterChannel(Channel): class TwitterChannel(Channel):
@@ -17,98 +18,56 @@ class TwitterChannel(Channel):
return "x.com" in d or "twitter.com" in d return "x.com" in d or "twitter.com" in d
def check(self, config=None): def check(self, config=None):
"""按 backends 顺序真实探测,第一个活着的后端即为 active_backend。""" # Prefer twitter-cli, fallback to bird/birdx
self.active_backend = None twitter = shutil.which("twitter")
failures = [] bird = shutil.which("bird") or shutil.which("birdx")
for backend in self.ordered_backends(config): if twitter:
if backend == "twitter-cli": return self._check_twitter_cli(twitter)
result = self._check_twitter_cli() elif bird:
elif backend == "bird CLI (legacy)": return self._check_bird(bird)
result = self._check_bird() else:
else:
continue
if result is None:
continue # 未安装——继续尝试下一个后端
status, message = result
if status in ("ok", "warn"):
# 工具本身是活的(含已装但未登录的 warn)
self.active_backend = backend
return status, message
# broken/timeout —— 记下处方,继续尝试下一个后端
failures.append(message)
if failures:
return "error", "\n".join(failures)
return "warn", (
"Twitter CLI 未安装。安装方式:\n"
" pipx install twitter-cli\n"
"或:\n"
" uv tool install twitter-cli"
)
def _check_twitter_cli(self):
"""探测 twitter-cli。返回 None 表示未安装,否则返回 (status, message)。
`twitter status` 才是健康信号:已登录时输出 "ok: true"
未登录时以非零退出码输出 "not_authenticated"——工具本身是活的,
所以 probe 的 error 状态也要看 output 内容再分类。
"""
probe = probe_command(
"twitter", ["status"], timeout=15, retries=1, package="twitter-cli"
)
if probe.status == "missing":
return None
if probe.status == "broken":
return "error", "twitter-cli 命令存在但无法执行。\n" + probe.hint
if probe.status == "timeout":
return "error", "twitter-cli 健康检查超时(已重试 1 次)。\n" + probe.hint
output = probe.output
if "ok: true" in output:
return "ok", (
"twitter-cli 完整可用(搜索、读推文、时间线、长文/Article、"
"用户查询、Thread"
)
if "not_authenticated" in output:
return "warn", ( return "warn", (
"twitter-cli 已安装但未认证。设置方式:\n" "Twitter CLI 未安装。安装方式:\n"
" export TWITTER_AUTH_TOKEN=\"xxx\"\n" " pipx install twitter-cli\n"
" export TWITTER_CT0=\"yyy\"\n" "或:\n"
"或确保已在浏览器中登录 x.com" " uv tool install twitter-cli"
) )
return "warn", (
"twitter-cli 已安装但认证检查失败。运行:\n"
" twitter -v status 查看详细信息"
)
def _check_bird(self): def _check_twitter_cli(self, binary: str):
"""探测 bird/birdxlegacy 回退)。返回 None 表示均未安装,否则返回 (status, message)。""" try:
last_failure = None r = subprocess.run(
for cmd in ("bird", "birdx"): [binary, "status"], capture_output=True,
probe = probe_command( encoding="utf-8", errors="replace", timeout=10
cmd, ["check"], timeout=15, retries=1, package="@steipete/bird"
) )
if probe.status == "missing": output = (r.stdout or "") + (r.stderr or "")
continue if r.returncode == 0 and "ok: true" in output:
if probe.status == "broken": return "ok", (
last_failure = ( "twitter-cli 完整可用(搜索、读推文、时间线、长文/Article、"
"error", "用户查询、Thread"
f"{cmd} 命令存在但无法执行(bird 是 npm 包,可用 "
"npm install -g @steipete/bird 重装)。\n" + probe.hint,
) )
continue # bird 坏了再试 birdx if "not_authenticated" in output:
if probe.status == "timeout": return "warn", (
last_failure = ( "twitter-cli 已安装但未认证。设置方式:\n"
"error", " export TWITTER_AUTH_TOKEN=\"xxx\"\n"
f"{cmd} 健康检查超时(已重试 1 次)。\n" + probe.hint, " export TWITTER_CT0=\"yyy\"\n"
"或确保已在浏览器中登录 x.com"
) )
continue return "warn", (
"twitter-cli 已安装但认证检查失败。运行:\n"
" twitter -v status 查看详细信息"
)
except Exception:
return "warn", "twitter-cli 已安装但连接失败"
output = probe.output def _check_bird(self, binary: str):
if probe.ok: try:
r = subprocess.run(
[binary, "check"], capture_output=True,
encoding="utf-8", errors="replace", timeout=10
)
output = (r.stdout or "") + (r.stderr or "")
if r.returncode == 0:
return "ok", "bird CLI 可用(读取、搜索推文,含长文/X Article)" return "ok", "bird CLI 可用(读取、搜索推文,含长文/X Article)"
if "Missing credentials" in output or "missing" in output.lower(): if "Missing credentials" in output or "missing" in output.lower():
return "warn", ( return "warn", (
@@ -119,4 +78,5 @@ class TwitterChannel(Channel):
return "warn", ( return "warn", (
"bird CLI 已安装但认证检查失败。" "bird CLI 已安装但认证检查失败。"
) )
return last_failure except Exception:
return "warn", "bird CLI 已安装但连接失败"
-2
View File
@@ -41,10 +41,8 @@ class V2EXChannel(Channel):
_get_json( _get_json(
"https://www.v2ex.com/api/topics/show.json?node_name=python&page=1" "https://www.v2ex.com/api/topics/show.json?node_name=python&page=1"
) )
self.active_backend = self.backends[0]
return "ok", "公开 API 可用(热门主题、节点浏览、主题详情、用户信息)" return "ok", "公开 API 可用(热门主题、节点浏览、主题详情、用户信息)"
except Exception as e: except Exception as e:
self.active_backend = None
return "warn", f"V2EX API 连接失败(可能需要代理):{e}" return "warn", f"V2EX API 连接失败(可能需要代理):{e}"
# ------------------------------------------------------------------ # # ------------------------------------------------------------------ #
-2
View File
@@ -17,8 +17,6 @@ class WebChannel(Channel):
return True # Fallback — handles any URL return True # Fallback — handles any URL
def check(self, config=None): def check(self, config=None):
# 恒可用兜底渠道:无本地命令、不做网络探测(doctor 已有多个渠道触网),保持零开销
self.active_backend = self.backends[0]
return "ok", "通过 Jina Reader 读取任意网页(curl https://r.jina.ai/URL" return "ok", "通过 Jina Reader 读取任意网页(curl https://r.jina.ai/URL"
def read(self, url: str) -> str: def read(self, url: str) -> str:
+63
View File
@@ -0,0 +1,63 @@
# -*- coding: utf-8 -*-
"""WeChat Official Account articles — read and search.
Read: Exa crawling (primary) / Camoufox stealth browser (optional)
Search: Exa web_search with includeDomains mp.weixin.qq.com
"""
import shutil
import subprocess
from .base import Channel
def _exa_available() -> bool:
mcporter = shutil.which("mcporter")
if not mcporter:
return False
try:
r = subprocess.run(
[mcporter, "config", "list"],
capture_output=True, encoding="utf-8", errors="replace", timeout=5,
)
return "exa" in r.stdout.lower()
except Exception:
return False
class WeChatChannel(Channel):
name = "wechat"
description = "微信公众号文章"
backends = ["Exa via mcporter (搜索+阅读)", "Camoufox (可选阅读)"]
tier = 0
def can_handle(self, url: str) -> bool:
from urllib.parse import urlparse
d = urlparse(url).netloc.lower()
return "mp.weixin.qq.com" in d or "weixin.qq.com" in d
def check(self, config=None):
has_exa = _exa_available()
has_camoufox = False
try:
import camoufox # noqa: F401
has_camoufox = True
except ImportError:
pass
if has_exa and has_camoufox:
return "ok", "完整可用(Exa 搜索 + Exa/Camoufox 阅读公众号文章)"
elif has_exa:
return "ok", (
"通过 Exa 搜索和阅读微信公众号文章(免费,无需额外配置)。"
"可选安装 Camoufox 获得更好的全文阅读效果。"
)
elif has_camoufox:
return "warn", (
"Camoufox 可阅读公众号文章,但搜索功能需要 Exa。"
"运行 `agent-reach install --env=auto` 安装 Exa。"
)
else:
return "off", (
"需要 mcporter + Exa MCP 来搜索和阅读微信公众号文章。\n"
"运行 `agent-reach install --env=auto` 安装。"
)
+52
View File
@@ -0,0 +1,52 @@
# -*- coding: utf-8 -*-
"""Weibo (微博) — check if mcporter + mcp-server-weibo is available."""
import shutil
import subprocess
from .base import Channel
class WeiboChannel(Channel):
name = "weibo"
description = "微博动态与热搜"
backends = ["mcp-server-weibo"]
tier = 1
def can_handle(self, url: str) -> bool:
from urllib.parse import urlparse
d = urlparse(url).netloc.lower()
return "weibo.com" in d or "weibo.cn" in d
def check(self, config=None):
mcporter = shutil.which("mcporter")
if not mcporter:
return "off", (
"需要 mcporter + mcp-server-weibo。安装步骤:\n"
" 1. npm install -g mcporter\n"
" 2. pip install git+https://github.com/Panniantong/mcp-server-weibo.git\n"
" 3. mcporter config add weibo --command 'mcp-server-weibo'\n"
" 详见 https://github.com/Panniantong/mcp-server-weibo"
)
try:
r = subprocess.run(
[mcporter, "config", "list"], capture_output=True,
encoding="utf-8", errors="replace", timeout=5
)
if "weibo" not in r.stdout:
return "off", (
"mcporter 已装但微博 MCP 未配置。运行:\n"
" pip install git+https://github.com/Panniantong/mcp-server-weibo.git\n"
" mcporter config add weibo --command 'mcp-server-weibo'"
)
except Exception:
return "off", "mcporter 连接异常"
try:
r = subprocess.run(
[mcporter, "list", "weibo"], capture_output=True,
encoding="utf-8", errors="replace", timeout=15
)
if r.returncode == 0 and "search_users" in r.stdout:
return "ok", "完整可用(热搜、搜索、用户动态、评论)"
return "warn", "MCP 已配置但工具加载失败,检查 mcp-server-weibo 版本"
except Exception:
return "warn", "MCP 连接异常,检查 mcp-server-weibo 是否可用"
+32 -128
View File
@@ -1,41 +1,10 @@
# -*- coding: utf-8 -*- # -*- coding: utf-8 -*-
"""XiaoHongShu — multi-backend: OpenCLI / xiaohongshu-mcp / xhs-cli. """XiaoHongShu — check if xhs-cli (xiaohongshu-cli) is available."""
Backend order encodes the recommendation, and probing order makes the
environment split automatic: OpenCLI needs a desktop Chrome so it simply
never probes alive on a server, where xiaohongshu-mcp (self-contained
headless browser) takes over. xhs-cli (upstream unmaintained since
2026-03) keeps working for existing installs as the last candidate.
"""
import urllib.error
import urllib.request
from agent_reach.probe import probe_command
import shutil
import subprocess
from .base import Channel from .base import Channel
_MCP_ENDPOINT = "http://localhost:18060/mcp"
_MCP_INSTALL_URL = "https://github.com/xpzouying/xiaohongshu-mcp"
def _mcp_service_reachable(timeout: int = 3) -> bool:
"""True if the xiaohongshu-mcp HTTP service answers on localhost.
Any HTTP response counts (the MCP endpoint replies 405 to GET) —
we only care that the service is up. Proxies are bypassed explicitly:
localhost must never be routed through HTTP_PROXY.
"""
req = urllib.request.Request(_MCP_ENDPOINT, method="GET")
opener = urllib.request.build_opener(urllib.request.ProxyHandler({}))
try:
opener.open(req, timeout=timeout)
return True
except urllib.error.HTTPError:
return True # 405/404 etc. — service is alive
except Exception:
return False
def format_xhs_result(data): def format_xhs_result(data):
"""Clean XHS API response, keeping only useful fields. """Clean XHS API response, keeping only useful fields.
@@ -149,7 +118,7 @@ def _clean_comment(comment):
class XiaoHongShuChannel(Channel): class XiaoHongShuChannel(Channel):
name = "xiaohongshu" name = "xiaohongshu"
description = "小红书笔记" description = "小红书笔记"
backends = ["OpenCLI", "xiaohongshu-mcp", "xhs-cli (xiaohongshu-cli)"] backends = ["xhs-cli (xiaohongshu-cli)"]
tier = 1 tier = 1
def can_handle(self, url: str) -> bool: def can_handle(self, url: str) -> bool:
@@ -158,101 +127,36 @@ class XiaoHongShuChannel(Channel):
return "xiaohongshu.com" in d or "xhslink.com" in d return "xiaohongshu.com" in d or "xhslink.com" in d
def check(self, config=None): def check(self, config=None):
"""Probe candidates in order; first fully-usable backend wins. xhs = shutil.which("xhs")
if not xhs:
If none is fully usable, the first fixable candidate (warn) is return "off", (
reported, so the user gets one actionable prescription instead "需要安装 xhs-cli\n"
of three half-relevant ones. " pipx install xiaohongshu-cli\n"
""" "或:\n"
self.active_backend = None " uv tool install xiaohongshu-cli\n"
findings = [] # (backend, status, message) "安装后运行 `xhs login` 登录"
for backend in self.ordered_backends(config):
if backend == "OpenCLI":
result = self._check_opencli()
elif backend == "xiaohongshu-mcp":
result = self._check_mcp()
else:
result = self._check_xhs_cli()
if result is None:
continue # not installed — not a candidate right now
findings.append((backend, *result))
for wanted in ("ok", "warn"):
for backend, status, message in findings:
if status == wanted:
self.active_backend = backend
return status, message
if findings: # only broken candidates left
return "error", "\n".join(m for _, _, m in findings)
return "off", (
"未安装任何小红书后端。推荐:\n"
" 桌面:agent-reach install --channels opencli\n"
" (复用 Chrome 登录态,刷过小红书即零配置可用)\n"
f" 服务器:xiaohongshu-mcp(自带无头浏览器+扫码登录):{_MCP_INSTALL_URL}"
)
def _check_opencli(self):
"""OpenCLI candidate. None = not installed."""
from agent_reach.backends import opencli_status
st = opencli_status()
if not st.installed:
return None
if st.broken:
return "error", st.hint
if st.ready:
return "ok", (
"OpenCLI 可用(复用浏览器登录态)。用法:"
"opencli xiaohongshu search/note/comments/feed -f yaml"
) )
return "warn", st.hint
def _check_mcp(self): try:
"""xiaohongshu-mcp candidate. None = service not running.""" r = subprocess.run(
if not _mcp_service_reachable(): [xhs, "status"], capture_output=True,
return None encoding="utf-8", errors="replace", timeout=10,
mcporter = probe_command(
"mcporter", ["config", "list"], timeout=10, package="mcporter"
)
if mcporter.ok and "xiaohongshu" in mcporter.output:
return "ok", (
"xiaohongshu-mcp 服务运行中"
"mcporter call 'xiaohongshu.search_feeds(keyword: \"...\")')。"
"若未登录,让 agent 调 get_login_qrcode 扫码"
) )
return "warn", ( output = (r.stdout or "") + (r.stderr or "")
"xiaohongshu-mcp 服务在跑但 mcporter 未接入。运行:\n" if r.returncode == 0 and "ok: true" in output:
f" mcporter config add xiaohongshu {_MCP_ENDPOINT}" return "ok", (
) "完整可用(搜索、阅读、评论、发帖、热门、"
"收藏、关注、用户查询)"
def _check_xhs_cli(self): )
"""Legacy xhs-cli candidate. None = not installed.""" if "not_authenticated" in output or "expired" in output:
probe = probe_command( return "warn", (
"xhs", ["status"], timeout=10, package="xiaohongshu-cli" "xhs-cli 已安装但未登录。运行:\n"
) " xhs login\n"
if probe.status == "missing": "(自动从浏览器提取 Cookie,或扫码登录)"
return None )
if probe.status == "broken":
return "error", "xhs 命令存在但无法执行\n" + probe.hint
if probe.status == "timeout":
return "warn", "xhs-cli 已安装但状态检测超时\n" + probe.hint
# 进程是活的(执行成功或运行后非零退出)——按输出内容分类
if probe.ok and "ok: true" in probe.output:
return "ok", (
"xhs-cli 可用(搜索、阅读、评论、热门;上游 2026-03 起停更,"
"桌面用户建议迁移到 OpenCLI"
)
if "not_authenticated" in probe.output or "expired" in probe.output:
return "warn", ( return "warn", (
"xhs-cli 已安装但未登录。运行:\n" "xhs-cli 已安装但状态异常。运行:\n"
" xhs login\n" " xhs -v status 查看详细信息"
"(自动从浏览器提取 Cookie,或扫码登录)"
) )
return "warn", ( except Exception:
"xhs-cli 已安装但状态异常。运行:\n" return "warn", "xhs-cli 已安装但连接失败"
" xhs -v status 查看详细信息"
)
+3 -12
View File
@@ -2,8 +2,8 @@
"""Xiaoyuzhou Podcast (小宇宙播客) — transcribe podcasts via Groq Whisper API.""" """Xiaoyuzhou Podcast (小宇宙播客) — transcribe podcasts via Groq Whisper API."""
import os import os
import shutil
from agent_reach.config import Config from agent_reach.config import Config
from agent_reach.probe import probe_command
from .base import Channel from .base import Channel
@@ -19,21 +19,13 @@ class XiaoyuzhouChannel(Channel):
return "xiaoyuzhoufm.com" in d return "xiaoyuzhoufm.com" in d
def check(self, config=None): def check(self, config=None):
self.active_backend = None # Check ffmpeg
if not shutil.which("ffmpeg"):
# Check ffmpeg — really execute it: a stale pip-installed ffmpeg shim
# passes shutil.which() but cannot run
probe = probe_command("ffmpeg", ["-version"], timeout=10, package="ffmpeg")
if probe.status == "missing":
return "off", ( return "off", (
"需要 ffmpeg(音频转码和切片)。安装:\n" "需要 ffmpeg(音频转码和切片)。安装:\n"
" Ubuntu/Debian: apt install -y ffmpeg\n" " Ubuntu/Debian: apt install -y ffmpeg\n"
" macOS: brew install ffmpeg" " macOS: brew install ffmpeg"
) )
if not probe.ok:
return "error", (
"ffmpeg 无法执行,重装:brew install ffmpegmacOS/ apt install ffmpegLinux"
)
# Check script exists # Check script exists
script = os.path.expanduser("~/.agent-reach/tools/xiaoyuzhou/transcribe.sh") script = os.path.expanduser("~/.agent-reach/tools/xiaoyuzhou/transcribe.sh")
@@ -59,5 +51,4 @@ class XiaoyuzhouChannel(Channel):
" 2. 运行: agent-reach configure groq-key gsk_xxxxx" " 2. 运行: agent-reach configure groq-key gsk_xxxxx"
) )
self.active_backend = "groq-whisper"
return "ok", "完整可用(播客下载 + Whisper 转录)" return "ok", "完整可用(播客下载 + Whisper 转录)"
-2
View File
@@ -162,14 +162,12 @@ class XueqiuChannel(Channel):
# ------------------------------------------------------------------ # # ------------------------------------------------------------------ #
def check(self, config=None): def check(self, config=None):
self.active_backend = None
try: try:
data = _get_json( data = _get_json(
"https://stock.xueqiu.com/v5/stock/batch/quote.json?symbol=SH000001" "https://stock.xueqiu.com/v5/stock/batch/quote.json?symbol=SH000001"
) )
items = (data.get("data") or {}).get("items") or [] items = (data.get("data") or {}).get("items") or []
if items: if items:
self.active_backend = self.backends[0]
return "ok", "公开 API 可用(行情、搜索、热帖、热股)" return "ok", "公开 API 可用(行情、搜索、热帖、热股)"
return "warn", "API 响应异常(返回数据为空)" return "warn", "API 响应异常(返回数据为空)"
except Exception as e: except Exception as e:
+8 -53
View File
@@ -3,23 +3,12 @@
import shutil import shutil
from agent_reach.probe import probe_command
from agent_reach.utils.paths import get_ytdlp_config_path, render_ytdlp_fix_command from agent_reach.utils.paths import get_ytdlp_config_path, render_ytdlp_fix_command
from agent_reach.utils.text import read_utf8_text from agent_reach.utils.text import read_utf8_text
from .base import Channel from .base import Channel
def _has_js_runtime_config(config_path) -> bool:
"""Return whether yt-dlp config explicitly enables a JS runtime."""
try:
if not config_path.exists():
return False
return "--js-runtimes" in read_utf8_text(config_path)
except OSError:
return False
class YouTubeChannel(Channel): class YouTubeChannel(Channel):
name = "youtube" name = "youtube"
description = "YouTube 视频和字幕" description = "YouTube 视频和字幕"
@@ -28,25 +17,12 @@ class YouTubeChannel(Channel):
def can_handle(self, url: str) -> bool: def can_handle(self, url: str) -> bool:
from urllib.parse import urlparse from urllib.parse import urlparse
d = urlparse(url).netloc.lower() d = urlparse(url).netloc.lower()
return "youtube.com" in d or "youtu.be" in d return "youtube.com" in d or "youtu.be" in d
def check(self, config=None): def check(self, config=None):
# 真跑 yt-dlp --version 探活,区分未装 / venv 断链 / 跑不动 if not shutil.which("yt-dlp"):
probe = probe_command("yt-dlp", ["--version"], timeout=10, package="yt-dlp")
if probe.status == "missing":
self.active_backend = None
return "off", "yt-dlp 未安装。安装:pip install yt-dlp" return "off", "yt-dlp 未安装。安装:pip install yt-dlp"
if probe.status == "broken":
self.active_backend = None
return "error", f"yt-dlp 已安装但无法执行\n{probe.hint}"
if not probe.ok: # timeout / error:装了但跑不动
self.active_backend = None
detail = probe.hint or probe.output or probe.status
return "error", f"yt-dlp 无法正常运行:{detail}"
# yt-dlp 本体是活的;后面的 JS runtime/转写检查只影响 ok/warn,不影响后端归属
self.active_backend = "yt-dlp"
# Check JS runtime # Check JS runtime
has_js = shutil.which("deno") or shutil.which("node") has_js = shutil.which("deno") or shutil.which("node")
if not has_js: if not has_js:
@@ -59,33 +35,12 @@ class YouTubeChannel(Channel):
has_deno = shutil.which("deno") has_deno = shutil.which("deno")
if not has_deno: if not has_deno:
ytdlp_config = get_ytdlp_config_path() ytdlp_config = get_ytdlp_config_path()
if not _has_js_runtime_config(ytdlp_config): has_js_config = False
if ytdlp_config.exists():
has_js_config = "--js-runtimes" in read_utf8_text(ytdlp_config)
if not has_js_config:
return "warn", ( return "warn", (
f"yt-dlp 已安装但未配置 JS runtime。运行:\n {render_ytdlp_fix_command()}" "yt-dlp 已安装但未配置 JS runtime。运行:\n"
f" {render_ytdlp_fix_command()}"
) )
# Surface transcription readiness so `doctor` reports it. return "ok", "可提取视频信息和字幕"
msg = "可提取视频信息和字幕"
if config is not None:
providers = []
if config.is_configured("groq_whisper"):
providers.append("groq")
if config.is_configured("openai_whisper"):
providers.append("openai")
if providers:
if not shutil.which("ffmpeg"):
msg += "(音频转写需安装 ffmpeg"
else:
msg += f",可转写音频({''.join(providers)}"
return "ok", msg
def transcribe(self, url: str, *, provider: str = "auto", config=None) -> str:
"""Download a YouTube video's audio and return its transcript.
Delegates to :func:`agent_reach.transcribe.transcribe`. Imported lazily
so the channel module stays cheap to import for users who never
transcribe.
"""
from agent_reach.transcribe import transcribe as _transcribe
return _transcribe(url, provider=provider, config=config)
+187 -175
View File
@@ -17,9 +17,6 @@ import time
from agent_reach import __version__ from agent_reach import __version__
# Pinned to the 0.4.2 state — PyPI still only has 0.4.1 (upstream issue #10).
_RDT_GIT_SOURCE = "git+https://github.com/public-clis/rdt-cli.git@5e4fb3720d5c174e976cd425ccc3b879d52cac66"
def _ensure_utf8_console(): def _ensure_utf8_console():
"""Best-effort Windows console UTF-8 setup for CLI runtime only.""" """Best-effort Windows console UTF-8 setup for CLI runtime only."""
@@ -73,13 +70,13 @@ def main():
help="Show what would be done without making any changes") help="Show what would be done without making any changes")
p_install.add_argument("--channels", default="", p_install.add_argument("--channels", default="",
help="Comma-separated optional channels to install " help="Comma-separated optional channels to install "
"(twitter,xiaoyuzhou,xueqiu,xiaohongshu," "(twitter,weibo,wechat,xiaoyuzhou,xueqiu,xiaohongshu,"
"reddit,bilibili,linkedin,all)") "reddit,bilibili,douyin,linkedin,all)")
# ── configure ── # ── configure ──
p_conf = sub.add_parser("configure", help="Set a config value or auto-extract from browser") p_conf = sub.add_parser("configure", help="Set a config value or auto-extract from browser")
p_conf.add_argument("key", nargs="?", default=None, p_conf.add_argument("key", nargs="?", default=None,
choices=["proxy", "github-token", "groq-key", "openai-key", choices=["proxy", "github-token", "groq-key",
"twitter-cookies", "youtube-cookies", "twitter-cookies", "youtube-cookies",
"xhs-cookies"], "xhs-cookies"],
help="What to configure (omit if using --from-browser)") help="What to configure (omit if using --from-browser)")
@@ -89,9 +86,7 @@ def main():
help="Auto-extract ALL platform cookies from browser (chrome/firefox/edge/brave/opera)") help="Auto-extract ALL platform cookies from browser (chrome/firefox/edge/brave/opera)")
# ── doctor ── # ── doctor ──
p_doctor = sub.add_parser("doctor", help="Check platform availability") sub.add_parser("doctor", help="Check platform availability")
p_doctor.add_argument("--json", action="store_true",
help="Output machine-readable JSON instead of the text report")
# ── uninstall ── # ── uninstall ──
p_uninstall = sub.add_parser("uninstall", help="Remove all Agent Reach config, tokens, and skill files") p_uninstall = sub.add_parser("uninstall", help="Remove all Agent Reach config, tokens, and skill files")
@@ -113,14 +108,6 @@ def main():
p_format.add_argument("platform", choices=["xhs"], help="Platform to format (xhs)") p_format.add_argument("platform", choices=["xhs"], help="Platform to format (xhs)")
# ── check-update ── # ── check-update ──
# ── transcribe ──
p_tr = sub.add_parser("transcribe", help="Transcribe a URL or local audio file (Whisper via Groq/OpenAI)")
p_tr.add_argument("source", help="Audio/video URL or local file path")
p_tr.add_argument("--provider", choices=["auto", "groq", "openai"], default="auto",
help="Transcription provider (default: auto = groq → openai fallback)")
p_tr.add_argument("-o", "--output", default=None,
help="Write transcript to a file instead of stdout")
sub.add_parser("check-update", help="Check for new versions and changes") sub.add_parser("check-update", help="Check for new versions and changes")
# ── watch ── # ── watch ──
@@ -143,7 +130,7 @@ def main():
sys.exit(0) sys.exit(0)
if args.command == "doctor": if args.command == "doctor":
_cmd_doctor(args) _cmd_doctor()
elif args.command == "check-update": elif args.command == "check-update":
_cmd_check_update() _cmd_check_update()
elif args.command == "watch": elif args.command == "watch":
@@ -160,8 +147,6 @@ def main():
_cmd_skill(args) _cmd_skill(args)
elif args.command == "format": elif args.command == "format":
_cmd_format(args) _cmd_format(args)
elif args.command == "transcribe":
_cmd_transcribe(args)
# ── Command handlers ──────────────────────────────── # ── Command handlers ────────────────────────────────
@@ -195,13 +180,14 @@ def _cmd_install(args):
# ── Parse --channels ── # ── Parse --channels ──
CHANNEL_INSTALLERS = { CHANNEL_INSTALLERS = {
"twitter": _install_twitter_deps, "twitter": _install_twitter_deps,
"weibo": _install_weibo_deps,
"wechat": _install_wechat_deps,
"xiaoyuzhou": _install_xiaoyuzhou_deps, "xiaoyuzhou": _install_xiaoyuzhou_deps,
"xiaohongshu": _install_xhs_deps, "xiaohongshu": _install_xhs_deps,
"reddit": _install_reddit_deps, "reddit": _install_reddit_deps,
"bilibili": _install_bili_deps, "bilibili": _install_bili_deps,
"opencli": _install_opencli_deps, # cross-channel backend, desktop only
# xueqiu: cookie-only, no install step # xueqiu: cookie-only, no install step
# linkedin: manual setup, no auto-install # douyin/linkedin: manual setup, no auto-install
} }
COOKIE_CHANNELS = {"twitter", "xueqiu", "bilibili"} COOKIE_CHANNELS = {"twitter", "xueqiu", "bilibili"}
@@ -209,7 +195,7 @@ def _cmd_install(args):
if args.channels: if args.channels:
raw = [c.strip().lower() for c in args.channels.split(",") if c.strip()] raw = [c.strip().lower() for c in args.channels.split(",") if c.strip()]
if "all" in raw: if "all" in raw:
requested_channels = set(CHANNEL_INSTALLERS.keys()) | {"xueqiu", "linkedin"} requested_channels = set(CHANNEL_INSTALLERS.keys()) | {"xueqiu", "douyin", "linkedin"}
else: else:
requested_channels = set(raw) requested_channels = set(raw)
@@ -253,10 +239,6 @@ def _cmd_install(args):
if requested_channels and not dry_run and not safe_mode: if requested_channels and not dry_run and not safe_mode:
print() print()
print("Installing optional channels...") print("Installing optional channels...")
if env == "server" and "opencli" in requested_channels:
# OpenCLI rides a real desktop Chrome session — useless headless
requested_channels.discard("opencli")
print(" -- OpenCLI 需要桌面环境 + Chrome,服务器环境跳过")
for ch_name in sorted(requested_channels): for ch_name in sorted(requested_channels):
installer = CHANNEL_INSTALLERS.get(ch_name) installer = CHANNEL_INSTALLERS.get(ch_name)
if installer: if installer:
@@ -325,7 +307,7 @@ def _cmd_install(args):
# First install — hint about optional channels # First install — hint about optional channels
print() print()
print("More channels available! Use --channels to install:") print("More channels available! Use --channels to install:")
print(" agent-reach install --channels=twitter,xiaohongshu,reddit,...") print(" agent-reach install --channels=twitter,weibo,xiaohongshu,...")
print(" agent-reach install --channels=all (install everything)") print(" agent-reach install --channels=all (install everything)")
# Star reminder # Star reminder
@@ -369,11 +351,8 @@ def _install_skill():
def _copy_skill_dir(target: str) -> bool: def _copy_skill_dir(target: str) -> bool:
"""Copy entire skill directory (locale-specific SKILL.md + references/).""" """Copy entire skill directory (locale-specific SKILL.md + references/)."""
try: try:
# Clear existing installation. A symlinked skill dir (dotfiles # Clear existing installation
# setups) breaks shutil.rmtree — unlink the link itself instead. if os.path.exists(target):
if os.path.islink(target):
os.unlink(target)
elif os.path.exists(target):
shutil.rmtree(target) shutil.rmtree(target)
os.makedirs(target, exist_ok=True) os.makedirs(target, exist_ok=True)
@@ -462,10 +441,7 @@ def _uninstall_skill():
skill_path = os.path.expanduser(skill_path_template) skill_path = os.path.expanduser(skill_path_template)
if os.path.isdir(skill_path): if os.path.isdir(skill_path):
try: try:
if os.path.islink(skill_path): shutil.rmtree(skill_path)
os.unlink(skill_path)
else:
shutil.rmtree(skill_path)
print(f" Removed {platform_name} skill: {skill_path}") print(f" Removed {platform_name} skill: {skill_path}")
removed = True removed = True
except Exception as e: except Exception as e:
@@ -627,7 +603,7 @@ def _install_system_deps():
except Exception: except Exception:
print(" -- Could not configure yt-dlp JS runtime (YouTube may not work)") print(" -- Could not configure yt-dlp JS runtime (YouTube may not work)")
# NOTE: twitter-cli, xiaoyuzhou, xhs-cli etc. are optional. # NOTE: twitter-cli, weibo, xiaoyuzhou, wechat, xhs-cli etc. are optional.
# They are installed via --channels flag, not here. # They are installed via --channels flag, not here.
# See CHANNEL_INSTALLERS in _cmd_install(). # See CHANNEL_INSTALLERS in _cmd_install().
@@ -699,78 +675,26 @@ def _install_twitter_deps():
def _install_xhs_deps(): def _install_xhs_deps():
"""Set up XiaoHongShu — backend depends on environment. """Install xhs-cli (xiaohongshu-cli) for XiaoHongShu."""
Desktop: OpenCLI (reuses the browser session, zero config).
Server: xiaohongshu-mcp guide (self-contained headless browser + QR
login; we don't manage long-running services, so guide only).
xhs-cli is no longer installed by default — upstream unmaintained
since 2026-03; existing installs keep working as a fallback backend.
"""
import shutil
print("Setting up XiaoHongShu...")
if _detect_environment() == "server":
print(" 服务器环境推荐 xiaohongshu-mcp(自带无头浏览器,扫码登录):")
print(" 1. 下载 binaryhttps://github.com/xpzouying/xiaohongshu-mcp/releases")
print(" (建议放到 ~/.agent-reach/tools/ 下)")
print(" 2. 启动服务(首次运行会下载约 150MB 浏览器,请等待完成)")
print(" 3. 扫码登录后接入:mcporter config add xiaohongshu http://localhost:18060/mcp")
print(" 4. 验证:agent-reach doctor")
return
_install_opencli_deps()
if shutil.which("xhs"):
print(" ✅ 检测到存量 xhs-cli,将作为备选后端继续可用")
def _install_opencli_deps():
"""Install OpenCLI — cross-platform backend riding the user's Chrome session.
Desktop-only. The npm package installs automatically; the Chrome
extension CANNOT be installed programmatically (Chrome security model),
so we print a one-click guide instead.
"""
import shutil import shutil
import subprocess import subprocess
from agent_reach.backends import ( print("Setting up XiaoHongShu (xhs-cli)...")
OPENCLI_EXTENSION_URL, if shutil.which("xhs"):
OPENCLI_PACKAGE, print(" ✅ xhs-cli already installed")
opencli_status,
opencli_summary,
)
print("Setting up OpenCLI (browser-session backend, desktop only)...")
st = opencli_status()
if st.installed and not st.broken:
print(f"{opencli_summary(st)}")
if not st.ready:
print(f" {st.hint}")
return return
for tool, cmd in [("pipx", ["pipx", "install", "xiaohongshu-cli"]),
if not shutil.which("npm"): ("uv", ["uv", "tool", "install", "xiaohongshu-cli"])]:
print(" [!] OpenCLI requires Node.js ≥ 20. Install Node first:") if shutil.which(tool):
print(" https://nodejs.org (或 brew install node") try:
return subprocess.run(cmd, capture_output=True, encoding="utf-8",
errors="replace", timeout=120)
try: if shutil.which("xhs"):
subprocess.run( print(" ✅ xhs-cli installed (run `xhs login` to authenticate)")
["npm", "install", "-g", OPENCLI_PACKAGE], return
capture_output=True, encoding="utf-8", errors="replace", timeout=300, except Exception:
) pass
except Exception: print(" [!] xhs-cli install failed. Run: pipx install xiaohongshu-cli")
pass
st = opencli_status()
if st.installed and not st.broken:
print(" ✅ OpenCLI installed")
print(" 最后一步(必须手动,Chrome 安全限制):安装浏览器扩展")
print(f" 1. 打开 {OPENCLI_EXTENSION_URL}")
print(" 2. 点「添加至 Chrome」")
print(" 3. 运行 `opencli doctor` 验证连接")
else:
print(f" [!] OpenCLI install failed. Run: npm install -g {OPENCLI_PACKAGE}")
def _install_reddit_deps(): def _install_reddit_deps():
@@ -782,10 +706,8 @@ def _install_reddit_deps():
if shutil.which("rdt"): if shutil.which("rdt"):
print(" ✅ rdt-cli already installed") print(" ✅ rdt-cli already installed")
return return
for tool, cmd in [ for tool, cmd in [("pipx", ["pipx", "install", "rdt-cli"]),
("pipx", ["pipx", "install", _RDT_GIT_SOURCE]), ("uv", ["uv", "tool", "install", "rdt-cli"])]:
("uv", ["uv", "tool", "install", "--from", _RDT_GIT_SOURCE, "rdt-cli"]),
]:
if shutil.which(tool): if shutil.which(tool):
try: try:
subprocess.run(cmd, capture_output=True, encoding="utf-8", subprocess.run(cmd, capture_output=True, encoding="utf-8",
@@ -795,7 +717,7 @@ def _install_reddit_deps():
return return
except Exception: except Exception:
pass pass
print(f" [!] rdt-cli install failed. Run: pipx install '{_RDT_GIT_SOURCE}'") print(" [!] rdt-cli install failed. Run: pipx install rdt-cli")
def _install_bili_deps(): def _install_bili_deps():
@@ -821,6 +743,123 @@ def _install_bili_deps():
print(" [!] bili-cli install failed. Run: pipx install bilibili-cli") print(" [!] bili-cli install failed. Run: pipx install bilibili-cli")
def _install_weibo_deps():
"""Install Weibo MCP server (Panniantong fork with visitor passport auth)."""
import shutil
import subprocess
print("Setting up Weibo MCP server...")
# Check if already installed and working
mcporter = shutil.which("mcporter")
if mcporter:
try:
r = subprocess.run(
[mcporter, "config", "list"], capture_output=True,
encoding="utf-8", errors="replace", timeout=5
)
if "weibo" in r.stdout:
print(" ✅ Weibo MCP already configured")
return
except Exception:
pass
# Install from our fork (has visitor passport auth fix)
try:
subprocess.run(
[sys.executable, "-m", "pip", "install", "-q",
"git+https://github.com/Panniantong/mcp-server-weibo.git"],
check=True, timeout=120
)
print(" ✅ mcp-server-weibo installed (Panniantong fork)")
except Exception as e:
print(f" [!] mcp-server-weibo install failed: {e}")
return
# Register with mcporter
if mcporter:
try:
subprocess.run(
[mcporter, "config", "add", "weibo", "--command", "mcp-server-weibo"],
check=True, capture_output=True, timeout=10
)
print(" ✅ Weibo MCP registered with mcporter")
except Exception:
print(" [!] mcporter config add failed. Run manually: mcporter config add weibo --command 'mcp-server-weibo'")
else:
print(" -- mcporter not found, skipping MCP registration. Install mcporter first, then run: mcporter config add weibo --command 'mcp-server-weibo'")
def _install_wechat_deps():
"""Install WeChat article reading and search dependencies."""
import subprocess
print("Setting up WeChat article tools...")
# Check if already installed
has_camoufox = False
has_miku = False
try:
import camoufox # noqa: F401
has_camoufox = True
except ImportError:
pass
try:
import miku_ai # noqa: F401
has_miku = True
except ImportError:
pass
# Install Python packages
if has_camoufox and has_miku:
print(" ✅ WeChat Python packages already installed")
else:
pkgs = []
if not has_camoufox:
pkgs.extend(["camoufox[geoip]", "markdownify", "beautifulsoup4", "httpx"])
if not has_miku:
pkgs.append("miku_ai")
try:
cmd = [sys.executable, "-m", "pip", "install", "--break-system-packages", "-q"] + pkgs
subprocess.run(cmd, capture_output=True, encoding="utf-8", errors="replace", timeout=120)
# Verify
ok = True
try:
import importlib
if not has_camoufox:
importlib.import_module("camoufox")
if not has_miku:
importlib.import_module("miku_ai")
except ImportError:
ok = False
if ok:
print(f" ✅ WeChat Python packages installed ({', '.join(pkgs)})")
else:
print(f" [!] Some WeChat packages failed to install. Try: pip install {' '.join(pkgs)}")
except Exception:
print(f" [!] WeChat packages install failed. Try: pip install {' '.join(pkgs)}")
# Clone wechat-article-for-ai tool
tools_dir = os.path.expanduser("~/.agent-reach/tools")
wechat_dir = os.path.join(tools_dir, "wechat-article-for-ai")
if os.path.isfile(os.path.join(wechat_dir, "main.py")):
print(" ✅ wechat-article-for-ai tool already installed")
else:
try:
os.makedirs(tools_dir, exist_ok=True)
subprocess.run(
["git", "clone", "--depth", "1",
"https://github.com/Panniantong/wechat-article-for-ai.git", wechat_dir],
capture_output=True, encoding="utf-8", errors="replace", timeout=60,
)
if os.path.isfile(os.path.join(wechat_dir, "main.py")):
print(" ✅ wechat-article-for-ai tool installed")
else:
print(" [!] wechat-article-for-ai clone failed. Try: git clone https://github.com/Panniantong/wechat-article-for-ai.git " + wechat_dir)
except Exception:
print(" [!] wechat-article-for-ai clone failed. Try: git clone https://github.com/Panniantong/wechat-article-for-ai.git " + wechat_dir)
def _install_system_deps_safe(): def _install_system_deps_safe():
"""Safe mode: check what's installed, print instructions for what's missing.""" """Safe mode: check what's installed, print instructions for what's missing."""
import shutil import shutil
@@ -1081,29 +1120,6 @@ def _cmd_configure(args):
config.set("groq_api_key", value) config.set("groq_api_key", value)
print(f"✅ Groq key configured!") print(f"✅ Groq key configured!")
elif args.key == "openai-key":
config.set("openai_api_key", value)
print(f"✅ OpenAI key configured!")
def _cmd_transcribe(args):
"""Transcribe a URL or local audio file via Whisper (Groq → OpenAI fallback)."""
from pathlib import Path
from agent_reach.transcribe import TranscribeError, transcribe
try:
text = transcribe(args.source, provider=args.provider)
except TranscribeError as e:
print(f"{e}")
sys.exit(1)
if args.output:
Path(args.output).write_text(text + "\n", encoding="utf-8")
print(f"✅ Transcript written to {args.output}")
else:
print(text)
def _parse_twitter_cookie_input(value: str): def _parse_twitter_cookie_input(value: str):
"""Parse Twitter cookie input from either separate values or a cookie header.""" """Parse Twitter cookie input from either separate values or a cookie header."""
@@ -1127,19 +1143,19 @@ def _parse_twitter_cookie_input(value: str):
def _configure_xhs_cookies(value): def _configure_xhs_cookies(value):
"""Import cookies into xiaohongshu-mcp Docker container. """Import cookies for xhs-cli, with legacy xiaohongshu-mcp support.
Accepts two formats: Accepts two formats:
1. Cookie-Editor JSON export (array of cookie objects) 1. Cookie-Editor JSON export (array of cookie objects)
2. Header String: "name1=value1; name2=value2; ..." 2. Header String: "name1=value1; name2=value2; ..."
The xiaohongshu-mcp container stores cookies at $COOKIES_PATH xhs-cli stores cookies as a name/value dict at ~/.xiaohongshu-cli/cookies.json.
(default: /app/data/cookies.json or cookies.json in workdir). The legacy Docker MCP stores Cookie-Editor style arrays at $COOKIES_PATH.
Format: JSON array of {name, value, domain, path, expires, httpOnly, secure, sameSite}.
""" """
import json import json
import shutil import shutil
import subprocess import subprocess
import time
value = value.strip() value = value.strip()
if not value: if not value:
@@ -1149,6 +1165,7 @@ def _configure_xhs_cookies(value):
# Detect format and parse # Detect format and parse
cookies_json = None cookies_json = None
cookies = []
# Try JSON format first (Cookie-Editor JSON export) # Try JSON format first (Cookie-Editor JSON export)
if value.startswith("["): if value.startswith("["):
@@ -1158,6 +1175,7 @@ def _configure_xhs_cookies(value):
# Validate it looks like cookie objects # Validate it looks like cookie objects
first = parsed[0] first = parsed[0]
if isinstance(first, dict) and "name" in first and "value" in first: if isinstance(first, dict) and "name" in first and "value" in first:
cookies = parsed
cookies_json = json.dumps(parsed) cookies_json = json.dumps(parsed)
print(f" Parsed {len(parsed)} cookies from JSON format") print(f" Parsed {len(parsed)} cookies from JSON format")
else: else:
@@ -1172,7 +1190,6 @@ def _configure_xhs_cookies(value):
# Header String format: "key1=val1; key2=val2; ..." # Header String format: "key1=val1; key2=val2; ..."
if cookies_json is None and "=" in value: if cookies_json is None and "=" in value:
cookies = []
for part in value.split(";"): for part in value.split(";"):
part = part.strip() part = part.strip()
if "=" not in part: if "=" not in part:
@@ -1206,34 +1223,37 @@ def _configure_xhs_cookies(value):
print(' 2. Header String: "key1=val1; key2=val2; ..."') print(' 2. Header String: "key1=val1; key2=val2; ..."')
return return
# Primary path: configure xhs-cli directly.
xhs_cookie_map = {
str(c.get("name", "")).strip(): str(c.get("value", ""))
for c in cookies
if isinstance(c, dict) and str(c.get("name", "")).strip()
}
if xhs_cookie_map.get("a1"):
xhs_config_dir = os.path.expanduser("~/.xiaohongshu-cli")
os.makedirs(xhs_config_dir, exist_ok=True)
xhs_cookie_path = os.path.join(xhs_config_dir, "cookies.json")
with open(xhs_cookie_path, "w", encoding="utf-8") as f:
json.dump({**xhs_cookie_map, "saved_at": time.time()}, f, indent=2)
os.chmod(xhs_cookie_path, 0o600)
print(f"✅ xhs-cli cookies saved to {xhs_cookie_path}")
print(" Run `xhs status` or `agent-reach doctor` to verify.")
else:
print("[!] Cookie input does not include the required `a1` cookie.")
print(" xhs-cli will not treat this as a logged-in session. Export all xiaohongshu.com cookies from Cookie-Editor.")
# Keep a legacy Cookie-Editor array for users still running xiaohongshu-mcp.
legacy_cookie_path = os.path.expanduser("~/.agent-reach/xhs-cookies.json")
os.makedirs(os.path.dirname(legacy_cookie_path), exist_ok=True)
with open(legacy_cookie_path, "w") as f:
f.write(cookies_json)
os.chmod(legacy_cookie_path, 0o600)
print(f" Legacy MCP cookie array saved to {legacy_cookie_path}")
# Find the container # Find the container
docker = shutil.which("docker") docker = shutil.which("docker")
if not docker: if not docker:
# No Docker - write to a local file for manual import. print(" Docker not found; skipping xiaohongshu-mcp import.")
# Create with 0o600 atomically so the file is never world-readable
# between open() and a follow-up chmod() (same pattern Config.save()
# uses in config.py).
import stat
cookie_path = os.path.expanduser("~/.agent-reach/xhs-cookies.json")
try:
fd = os.open(
cookie_path,
os.O_WRONLY | os.O_CREAT | os.O_TRUNC,
stat.S_IRUSR | stat.S_IWUSR, # 0o600
)
with os.fdopen(fd, "w", encoding="utf-8") as f:
f.write(cookies_json)
except OSError:
# Windows / unsupported flags — fall back to plain open + chmod.
with open(cookie_path, "w", encoding="utf-8") as f:
f.write(cookies_json)
try:
os.chmod(cookie_path, 0o600)
except OSError:
pass
print(f" Cookies saved to {cookie_path}")
print(" Docker not found. Copy manually:")
print(f" docker cp {cookie_path} xiaohongshu-mcp:/app/data/cookies.json")
return return
# Check if xiaohongshu-mcp container is running # Check if xiaohongshu-mcp container is running
@@ -1244,9 +1264,7 @@ def _configure_xhs_cookies(value):
) )
container_name = result.stdout.strip() container_name = result.stdout.strip()
if not container_name: if not container_name:
print("[X] xiaohongshu-mcp container is not running.") print(" xiaohongshu-mcp container is not running; skipping legacy Docker import.")
print(" Start it first:")
print(" docker run -d --name xiaohongshu-mcp -p 18060:18060 xpzouying/xiaohongshu-mcp")
return return
except Exception as e: except Exception as e:
print(f"[X] Could not check Docker: {e}") print(f"[X] Could not check Docker: {e}")
@@ -1417,7 +1435,7 @@ def _cmd_uninstall(args):
print(" npm uninstall -g undici") print(" npm uninstall -g undici")
def _cmd_doctor(args=None): def _cmd_doctor():
from agent_reach.config import Config from agent_reach.config import Config
from agent_reach.doctor import check_all, format_report from agent_reach.doctor import check_all, format_report
try: try:
@@ -1426,11 +1444,6 @@ def _cmd_doctor(args=None):
rprint = print rprint = print
config = Config() config = Config()
results = check_all(config) results = check_all(config)
if args is not None and getattr(args, "json", False):
print(json.dumps(results, ensure_ascii=False, indent=2))
return
rprint(format_report(results)) rprint(format_report(results))
# Auto-install skill if not already present (fixes #154) # Auto-install skill if not already present (fixes #154)
@@ -1501,8 +1514,7 @@ def _cmd_setup():
# Step 3: Reddit — rdt-cli # Step 3: Reddit — rdt-cli
print("【信息】Reddit — 通过 rdt-cli 搜索和阅读,无需配置") print("【信息】Reddit — 通过 rdt-cli 搜索和阅读,无需配置")
print(f" 安装:pipx install '{_RDT_GIT_SOURCE}'") print(" 安装:pipx install rdt-cli")
print(" 然后运行:rdt login")
print() print()
# Step 4: Groq (Whisper) # Step 4: Groq (Whisper)
+1 -2
View File
@@ -21,9 +21,8 @@ class Config:
# Feature → required config keys # Feature → required config keys
FEATURE_REQUIREMENTS = { FEATURE_REQUIREMENTS = {
"exa_search": ["exa_api_key"], "exa_search": ["exa_api_key"],
"twitter_xreach": ["twitter_auth_token", "twitter_ct0"], # legacy key name; used by twitter-cli "twitter_xreach": ["twitter_auth_token", "twitter_ct0"], # legacy key name; used by bird CLI
"groq_whisper": ["groq_api_key"], "groq_whisper": ["groq_api_key"],
"openai_whisper": ["openai_api_key"],
"github_token": ["github_token"], "github_token": ["github_token"],
} }
+6 -29
View File
@@ -148,28 +148,6 @@ def extract_all(browser: str = "chrome") -> Dict[str, dict]:
return results return results
def _open_owner_only(path: str):
"""Open *path* for writing, atomically creating it with mode 0o600.
Mirrors the pattern used by Config.save() in config.py: O_WRONLY|O_CREAT|
O_TRUNC + an explicit mode argument so the file is never briefly
world-readable between open() and a later os.chmod(). On Windows (or any
OS that rejects the open flags) we fall back to a plain open().
"""
import os
import stat
try:
fd = os.open(
path,
os.O_WRONLY | os.O_CREAT | os.O_TRUNC,
stat.S_IRUSR | stat.S_IWUSR, # 0o600
)
return os.fdopen(fd, "w", encoding="utf-8")
except OSError:
return open(path, "w", encoding="utf-8")
def _sync_xfetch_session(auth_token: str, ct0: str) -> None: def _sync_xfetch_session(auth_token: str, ct0: str) -> None:
"""Sync Twitter credentials to ~/.config/xfetch/session.json (legacy xreach compat).""" """Sync Twitter credentials to ~/.config/xfetch/session.json (legacy xreach compat)."""
import json import json
@@ -188,8 +166,9 @@ def _sync_xfetch_session(auth_token: str, ct0: str) -> None:
session_data = {} session_data = {}
session_data["authToken"] = auth_token session_data["authToken"] = auth_token
session_data["ct0"] = ct0 session_data["ct0"] = ct0
with _open_owner_only(session_path) as sf: with open(session_path, "w", encoding="utf-8") as sf:
json.dump(session_data, sf, indent=2) json.dump(session_data, sf, indent=2)
os.chmod(session_path, 0o600)
except Exception: except Exception:
# Non-fatal: agent-reach config is the source of truth, xfetch sync is best-effort # Non-fatal: agent-reach config is the source of truth, xfetch sync is best-effort
pass pass
@@ -200,19 +179,17 @@ def _sync_bird_env(auth_token: str, ct0: str) -> None:
bird reads AUTH_TOKEN and CT0 from environment variables. This writes a bird reads AUTH_TOKEN and CT0 from environment variables. This writes a
shell-sourceable file so users can `source ~/.config/bird/credentials.env`. shell-sourceable file so users can `source ~/.config/bird/credentials.env`.
Values are passed through shlex.quote so a token containing a quote, $, or
backtick cannot break out into shell syntax when the file is sourced.
""" """
import os import os
import shlex
try: try:
bird_dir = os.path.join(os.path.expanduser("~"), ".config", "bird") bird_dir = os.path.join(os.path.expanduser("~"), ".config", "bird")
os.makedirs(bird_dir, exist_ok=True) os.makedirs(bird_dir, exist_ok=True)
env_path = os.path.join(bird_dir, "credentials.env") env_path = os.path.join(bird_dir, "credentials.env")
with _open_owner_only(env_path) as f: with open(env_path, "w", encoding="utf-8") as f:
f.write(f"AUTH_TOKEN={shlex.quote(auth_token)}\n") f.write(f'AUTH_TOKEN="{auth_token}"\n')
f.write(f"CT0={shlex.quote(ct0)}\n") f.write(f'CT0="{ct0}"\n')
os.chmod(env_path, 0o600)
except Exception: except Exception:
# Non-fatal: agent-reach config is the source of truth, bird env sync is best-effort # Non-fatal: agent-reach config is the source of truth, bird env sync is best-effort
pass pass
+7 -23
View File
@@ -10,37 +10,20 @@ from agent_reach.channels import get_all_channels
def check_all(config: Config) -> Dict[str, dict]: def check_all(config: Config) -> Dict[str, dict]:
"""Check all channels and return status dict. """Check all channels and return status dict."""
A single misbehaving channel must never take the whole report down,
so per-channel exceptions degrade to status="error".
"""
results = {} results = {}
for ch in get_all_channels(): for ch in get_all_channels():
try: status, message = ch.check(config)
status, message = ch.check(config)
except Exception as e: # noqa: BLE001 — doctor must survive any channel
status, message = "error", f"体检异常:{e}"
results[ch.name] = { results[ch.name] = {
"status": status, "status": status,
"name": ch.description, "name": ch.description,
"message": message, "message": message,
"tier": ch.tier, "tier": ch.tier,
"backends": ch.backends, "backends": ch.backends,
"active_backend": getattr(ch, "active_backend", None),
} }
return results return results
def _name_msg(r: dict, escape) -> str:
"""Render one channel line; show the active backend when there is a choice."""
text = f"[bold]{escape(r['name'])}[/bold] — {escape(r['message'])}"
active = r.get("active_backend")
if active and len(r.get("backends", [])) > 1:
text += f" [dim](当前后端:{escape(active)}[/dim]"
return text
def format_report(results: Dict[str, dict]) -> str: def format_report(results: Dict[str, dict]) -> str:
"""Format results as a readable text report (with Rich markup).""" """Format results as a readable text report (with Rich markup)."""
try: try:
@@ -51,7 +34,6 @@ def format_report(results: Dict[str, dict]) -> str:
lines = [] lines = []
lines.append("[bold cyan]Agent Reach 状态[/bold cyan]") lines.append("[bold cyan]Agent Reach 状态[/bold cyan]")
lines.append("[cyan]" + "=" * 40 + "[/cyan]") lines.append("[cyan]" + "=" * 40 + "[/cyan]")
lines.append("图例:[green]✅[/green] 可用 [yellow][!][/yellow] 已装但需配置/登录 [red][X][/red] 未安装")
ok_count = sum(1 for r in results.values() if r["status"] == "ok") ok_count = sum(1 for r in results.values() if r["status"] == "ok")
total = len(results) total = len(results)
@@ -61,7 +43,7 @@ def format_report(results: Dict[str, dict]) -> str:
lines.append("[bold]✅ 装好即用:[/bold]") lines.append("[bold]✅ 装好即用:[/bold]")
for key, r in results.items(): for key, r in results.items():
if r["tier"] == 0: if r["tier"] == 0:
name_msg = _name_msg(r, escape) name_msg = f"[bold]{escape(r['name'])}[/bold] — {escape(r['message'])}"
if r["status"] == "ok": if r["status"] == "ok":
lines.append(f" [green]✅[/green] {name_msg}") lines.append(f" [green]✅[/green] {name_msg}")
elif r["status"] == "warn": elif r["status"] == "warn":
@@ -77,7 +59,8 @@ def format_report(results: Dict[str, dict]) -> str:
lines.append("") lines.append("")
lines.append("[bold]可选渠道(已安装):[/bold]") lines.append("[bold]可选渠道(已安装):[/bold]")
for key, r in tier1_active.items(): for key, r in tier1_active.items():
lines.append(f" [green]✅[/green] {_name_msg(r, escape)}") name_msg = f"[bold]{escape(r['name'])}[/bold] — {escape(r['message'])}"
lines.append(f" [green]✅[/green] {name_msg}")
# Tier 2 — optional complex setup # Tier 2 — optional complex setup
tier2 = {k: r for k, r in results.items() if r["tier"] == 2} tier2 = {k: r for k, r in results.items() if r["tier"] == 2}
@@ -88,7 +71,8 @@ def format_report(results: Dict[str, dict]) -> str:
lines.append("") lines.append("")
lines.append("[bold]可选渠道(已安装):[/bold]") lines.append("[bold]可选渠道(已安装):[/bold]")
for key, r in tier2_active.items(): for key, r in tier2_active.items():
lines.append(f" [green]✅[/green] {_name_msg(r, escape)}") name_msg = f"[bold]{escape(r['name'])}[/bold] — {escape(r['message'])}"
lines.append(f" [green]✅[/green] {name_msg}")
lines.append("") lines.append("")
status_color = "green" if ok_count == total else ("yellow" if ok_count > 0 else "red") status_color = "green" if ok_count == total else ("yellow" if ok_count > 0 else "red")
+3 -3
View File
@@ -8,7 +8,7 @@ Agent Reach 通过 **rdt-cli** 实现 Reddit 的搜索和阅读功能:
- **搜索**`rdt search "关键词"` - **搜索**`rdt search "关键词"`
- **阅读完整帖子+评论**`rdt read POST_ID` - **阅读完整帖子+评论**`rdt read POST_ID`
免费,无需代理,无需 API Key。需要登录认证(`rdt login`,自动从浏览器提取 Cookie 免费,无需代理,无需 API Key,无需登录
## Agent 可自动完成的步骤 ## Agent 可自动完成的步骤
@@ -17,9 +17,9 @@ Agent Reach 通过 **rdt-cli** 实现 Reddit 的搜索和阅读功能:
which rdt && echo "installed" || echo "not installed" which rdt && echo "installed" || echo "not installed"
``` ```
2. 如果未安装,自动安装(PyPI 版本暂时落后,从 GitHub 安装最新版) 2. 如果未安装,自动安装:
```bash ```bash
pipx install 'git+https://github.com/public-clis/rdt-cli.git' pipx install rdt-cli
``` ```
或一键安装: 或一键安装:
+40
View File
@@ -0,0 +1,40 @@
# 微信公众号配置指南
## 功能说明
读取微信公众号文章。需要 Playwright 来处理微信的反爬机制。
## Agent 可自动完成的步骤
1. 检查 Playwright 是否安装:
```bash
python3 -c "import playwright; print('installed')" 2>&1
```
2. 安装 Playwright + 浏览器:
```bash
pip install playwright
playwright install chromium
```
3. 安装完成后测试:
```bash
curl -s "https://r.jina.ai/https://mp.weixin.qq.com/s/一个测试链接" -H "Accept: text/markdown"
```
## 需要用户手动做的步骤
请告诉用户:
> 微信公众号的配置很简单,只需要安装一个浏览器组件(约 150MB)。
>
> 我来帮你安装,你不需要做任何事情。安装过程大约 1-2 分钟。
>
> 安装好之后就可以直接读取微信公众号文章了,不需要登录。
## Agent 操作流程
1. 安装 Playwright`pip install playwright`
2. 安装 Chromium`playwright install chromium`
3. 测试:读一篇微信文章
4. 反馈:"✅ 微信公众号已配置!发给我任何公众号文章链接,我都能读取。"
5. 如果安装失败(空间不足等):"❌ 浏览器组件安装失败。可能是磁盘空间不足(需要约 150MB)。"
+6 -1
View File
@@ -39,7 +39,9 @@ agent-reach doctor
> 3. 点击 Cookie-Editor 图标 → Export → Header String > 3. 点击 Cookie-Editor 图标 → Export → Header String
> 4. 把导出的字符串发给 Agent,运行:`agent-reach configure xhs-cookies "导出的cookie字符串"` > 4. 把导出的字符串发给 Agent,运行:`agent-reach configure xhs-cookies "导出的cookie字符串"`
> >
> **注意**:不要依赖 QR 扫码登录,Cookie-Editor 导出方式最简单可靠。 > **注意**:不要依赖 QR 扫码登录,Cookie-Editor 导出方式最简单可靠。这个方式也适合 WSL、SSH、容器等无法直接读取桌面浏览器 Cookie 的环境。
`agent-reach configure xhs-cookies` 会把 Cookie 同步到 `~/.xiaohongshu-cli/cookies.json`,供 `xhs status/search/read` 直接使用;如果你还在用旧的 `xiaohongshu-mcp` Docker 方案,也会保留兼容导入文件。
## 使用示例 ## 使用示例
@@ -69,6 +71,9 @@ A: 推荐使用住宅代理:`export HTTP_PROXY="http://user:pass@ip:port"`。
**Q: xhs-cli 不支持我的系统?** **Q: xhs-cli 不支持我的系统?**
A: 确保 Python 3.10+ 和 pipx 已安装。运行 `pipx install xiaohongshu-cli` 即可。 A: 确保 Python 3.10+ 和 pipx 已安装。运行 `pipx install xiaohongshu-cli` 即可。
**Q: WSL 里 `xhs login` 读不到 Windows 浏览器怎么办?**
A: 这是常见限制。WSL 里没有直接可读的 Linux 浏览器 Cookie 时,`xhs login` 自动提取会失败。推荐在 Windows 浏览器用 Cookie-Editor 导出 `xiaohongshu.com` 的 Header String,然后在 WSL 里运行 `agent-reach configure xhs-cookies "..."`,再用 `xhs status` 验证。
## 备选方案:Docker MCP ## 备选方案:Docker MCP
如果你已经在使用 [xiaohongshu-mcp](https://github.com/xpzouying/xiaohongshu-mcp) Docker 方案,它也能正常工作: 如果你已经在使用 [xiaohongshu-mcp](https://github.com/xpzouying/xiaohongshu-mcp) Docker 方案,它也能正常工作:
-99
View File
@@ -1,99 +0,0 @@
# -*- coding: utf-8 -*-
"""Lightweight upstream command probing.
Distinguishes the three failure modes that look identical to shutil.which():
- missing: command not on PATH
- broken: command exists but cannot execute most commonly a stale venv
shebang after a system Python upgrade (pipx/uv tool installs break this
way: which() finds the shim, but exec fails with FileNotFoundError
pointing at the shim itself)
- timeout/error: command runs but misbehaves
Channels use probe_command() inside check() so doctor reports real health,
not just file existence.
"""
import shutil
import subprocess
from dataclasses import dataclass
from typing import Optional, Sequence
from agent_reach.utils.process import utf8_subprocess_env
#: Exit codes shells use for "found but not executable" / "not found".
_BROKEN_EXIT_CODES = (126, 127)
@dataclass
class ProbeResult:
status: str # "ok" | "missing" | "broken" | "timeout" | "error"
output: str = ""
hint: str = ""
@property
def ok(self) -> bool:
return self.status == "ok"
def reinstall_hint(package: str) -> str:
"""Prescription for a broken (stale-venv) CLI install."""
return (
f"命令存在但无法执行——通常是系统 Python 升级后 venv 解释器丢失。重装即可修复:\n"
f" uv tool install --force {package}\n"
f"或:pipx reinstall {package}"
)
def probe_command(
cmd: str,
args: Sequence[str] = ("--version",),
timeout: int = 10,
retries: int = 0,
package: Optional[str] = None,
) -> ProbeResult:
"""Actually execute `cmd *args` and classify the result.
package: pip/pipx package name used in the broken-install hint
(defaults to cmd).
"""
path = shutil.which(cmd)
if not path:
return ProbeResult("missing")
last: Optional[ProbeResult] = None
for _ in range(retries + 1):
last = _run_once(path, args, timeout, package or cmd)
if last.ok:
return last
# missing/broken won't heal between retries — only transient
# failures (timeout/error) are worth a second attempt
if last.status in ("missing", "broken"):
return last
return last
def _run_once(path: str, args: Sequence[str], timeout: int, package: str) -> ProbeResult:
try:
r = subprocess.run(
[path, *args],
capture_output=True,
encoding="utf-8",
errors="replace",
timeout=timeout,
env=utf8_subprocess_env(),
)
except FileNotFoundError:
# which() found it but exec failed: the shebang interpreter is gone
return ProbeResult("broken", hint=reinstall_hint(package))
except OSError:
return ProbeResult("broken", hint=reinstall_hint(package))
except subprocess.TimeoutExpired:
return ProbeResult("timeout", hint=f"`{path}` 响应超时(>{timeout}s")
if r.returncode in _BROKEN_EXIT_CODES:
return ProbeResult("broken", hint=reinstall_hint(package))
output = (r.stdout or "") + (r.stderr or "")
if r.returncode != 0:
return ProbeResult("error", output=output.strip())
return ProbeResult("ok", output=output.strip())
+7 -109
View File
@@ -1,29 +1,11 @@
#!/bin/bash #!/bin/bash
# 小宇宙播客转文字脚本 # 小宇宙播客转文字脚本
# 用法: bash transcribe.sh [--polish] <小宇宙链接> [输出文件路径] # 用法: bash transcribe.sh <小宇宙链接> [输出文件路径]
# 环境变量: GROQ_API_KEY (必须) # 环境变量: GROQ_API_KEY (必须)
#
# --polish: 转录后调用 Groq Llama 3.3 70B 给文稿补中文标点+合理分段
# (Whisper 对中文标点支持较弱,开启后阅读体验显著更好)
set -e set -e
POLISH=0 URL="${1:?用法: bash transcribe.sh <小宇宙链接> [输出文件路径]}"
while [ $# -gt 0 ]; do
case "$1" in
--polish) POLISH=1; shift ;;
--) shift; break ;;
-h|--help)
echo "用法: bash transcribe.sh [--polish] <小宇宙链接> [输出文件路径]"
exit 0 ;;
--*)
echo "未知选项: $1" >&2
exit 1 ;;
*) break ;;
esac
done
URL="${1:?用法: bash transcribe.sh [--polish] <小宇宙链接> [输出文件路径]}"
OUTPUT="${2:-/tmp/podcast_transcript.txt}" OUTPUT="${2:-/tmp/podcast_transcript.txt}"
TMPDIR="/tmp/xiaoyuzhou_$$" TMPDIR="/tmp/xiaoyuzhou_$$"
@@ -53,8 +35,8 @@ echo "===================="
# Step 1: 提取音频 URL 和标题 # Step 1: 提取音频 URL 和标题
echo "🔍 正在解析页面..." echo "🔍 正在解析页面..."
PAGE=$(curl -s "$URL") PAGE=$(curl -s "$URL")
AUDIO_URL=$(echo "$PAGE" | perl -ne 'while (/(https:\/\/media\.xyzcdn\.net\/[^"]*\.(?:m4a|mp3))/gi) { print "$1\n" }' | head -1) AUDIO_URL=$(echo "$PAGE" | grep -oP 'https://media\.xyzcdn\.net/[^"]*\.(m4a|mp3)' | head -1)
TITLE=$(echo "$PAGE" | perl -ne 'if (/"title":"([^"]*)"/) { print "$1\n"; last }' | head -1) TITLE=$(echo "$PAGE" | grep -oP '"title":"[^"]*"' | head -1 | sed 's/"title":"//;s/"//')
if [ -z "$AUDIO_URL" ]; then if [ -z "$AUDIO_URL" ]; then
echo "❌ 无法从页面提取音频链接" echo "❌ 无法从页面提取音频链接"
@@ -117,7 +99,6 @@ for i in $(seq 0 $((NUM_CHUNKS - 1))); do
-F file="@$TMPDIR/chunk_${i}.mp3" \ -F file="@$TMPDIR/chunk_${i}.mp3" \
-F model="whisper-large-v3" \ -F model="whisper-large-v3" \
-F language="zh" \ -F language="zh" \
-F prompt="以下是一段中文普通话播客录音,请输出包含完整中文标点(,。?!:;“”‘’)的转写文本。" \
-F response_format="text") -F response_format="text")
HTTP_CODE=$(echo "$RESPONSE" | tail -1) HTTP_CODE=$(echo "$RESPONSE" | tail -1)
@@ -130,7 +111,7 @@ for i in $(seq 0 $((NUM_CHUNKS - 1))); do
# 如果是速率限制,等待后重试 # 如果是速率限制,等待后重试
if [ "$HTTP_CODE" = "429" ]; then if [ "$HTTP_CODE" = "429" ]; then
# 从错误信息中提取等待时间,默认 120 秒 # 从错误信息中提取等待时间,默认 120 秒
WAIT_SEC=$(echo "$BODY" | perl -ne 'if (/in (\d+)m/) { print "$1\n"; exit }') WAIT_SEC=$(echo "$BODY" | grep -oP 'in \K[0-9]+m' | sed 's/m//' | head -1)
WAIT_SEC=${WAIT_SEC:-2} WAIT_SEC=${WAIT_SEC:-2}
WAIT_SEC=$((WAIT_SEC * 60 + 30)) WAIT_SEC=$((WAIT_SEC * 60 + 30))
echo " ⏳ 速率限制,等待 ${WAIT_SEC} 秒后重试..." echo " ⏳ 速率限制,等待 ${WAIT_SEC} 秒后重试..."
@@ -141,7 +122,6 @@ for i in $(seq 0 $((NUM_CHUNKS - 1))); do
-F file="@$TMPDIR/chunk_${i}.mp3" \ -F file="@$TMPDIR/chunk_${i}.mp3" \
-F model="whisper-large-v3" \ -F model="whisper-large-v3" \
-F language="zh" \ -F language="zh" \
-F prompt="以下是一段中文普通话播客录音,请输出包含完整中文标点(,。?!:;“”‘’)的转写文本。" \
-F response_format="text") -F response_format="text")
HTTP_CODE=$(echo "$RESPONSE" | tail -1) HTTP_CODE=$(echo "$RESPONSE" | tail -1)
BODY=$(echo "$RESPONSE" | sed '$d') BODY=$(echo "$RESPONSE" | sed '$d')
@@ -160,81 +140,6 @@ for i in $(seq 0 $((NUM_CHUNKS - 1))); do
echo "✅ ($CHARS 字)" echo "✅ ($CHARS 字)"
done done
# Step 6.5 (可选): 用 Llama 3.3 70B 给文稿补标点+分段
if [ "$POLISH" = "1" ]; then
echo "✨ 正在润色(Llama 3.3 70B 加标点+分段)..."
for i in $(seq 0 $((NUM_CHUNKS - 1))); do
echo -n "$((i+1))/$NUM_CHUNKS... "
IN_FILE="$TMPDIR/transcript_${i}.txt" \
OUT_FILE="$TMPDIR/polished_${i}.txt" \
GROQ_API_KEY="$GROQ_API_KEY" \
python3 <<'PY'
import json, os, sys, urllib.request, urllib.error
KEY = os.environ["GROQ_API_KEY"]
IN = os.environ["IN_FILE"]
OUT = os.environ["OUT_FILE"]
MODEL = "llama-3.3-70b-versatile"
MAX_DEPTH = 3
PROMPT_TMPL = (
"以下是一段中文普通话播客的语音转写片段,由于 Whisper 对中文标点支持较弱,"
"整段几乎没有标点。请你**只做一件事**:在合适位置补充中文标点(,。!?:;),"
"可以适度分段。\n\n"
"**严格要求**\n"
"- 不得修改、删除、增加任何汉字或英文/数字\n"
"- 不得改写、润色、总结\n"
"- 不得添加任何解释、前言、后记\n"
"- 直接输出加好标点+合理分段后的全文\n\n"
"原文:\n{}"
)
def call_groq(text):
body = json.dumps({
"model": MODEL,
"temperature": 0.2,
"max_completion_tokens": 8192,
"messages": [{"role": "user", "content": PROMPT_TMPL.format(text)}],
}).encode()
req = urllib.request.Request(
"https://api.groq.com/openai/v1/chat/completions",
data=body,
headers={
"Authorization": f"Bearer {KEY}",
"Content-Type": "application/json",
"User-Agent": "agent-reach-xiaoyuzhou/1.0",
},
)
with urllib.request.urlopen(req, timeout=180) as r:
resp = json.load(r)
return (
resp["choices"][0]["message"]["content"].strip(),
resp["choices"][0].get("finish_reason"),
)
def polish(text, depth=0):
try:
out, fr = call_groq(text)
except urllib.error.HTTPError as e:
sys.stderr.write(f"polish HTTP {e.code}: {e.read().decode(errors='replace')[:200]}\n")
return text # fallback to raw
except Exception as e:
sys.stderr.write(f"polish error: {e}\n")
return text
if fr != "length" or depth >= MAX_DEPTH:
return out
# 输出被截断:从中点切两半递归处理
mid = len(text) // 2
return polish(text[:mid], depth + 1) + polish(text[mid:], depth + 1)
content = open(IN, encoding="utf-8").read().strip()
result = polish(content)
open(OUT, "w", encoding="utf-8").write(result + "\n")
print(f"✅ ({len(result)} 字)")
PY
done
fi
# Step 7: 合并输出 # Step 7: 合并输出
echo "📄 正在合并文字稿..." echo "📄 正在合并文字稿..."
@@ -244,19 +149,12 @@ echo "📄 正在合并文字稿..."
echo "来源: $URL" echo "来源: $URL"
echo "时长: ${DURATION_MIN}${DURATION_SEC}" echo "时长: ${DURATION_MIN}${DURATION_SEC}"
echo "转录时间: $(date '+%Y-%m-%d %H:%M')" echo "转录时间: $(date '+%Y-%m-%d %H:%M')"
if [ "$POLISH" = "1" ]; then
echo "润色: Groq Llama 3.3 70B"
fi
echo "" echo ""
echo "---" echo "---"
echo "" echo ""
for i in $(seq 0 $((NUM_CHUNKS - 1))); do for i in $(seq 0 $((NUM_CHUNKS - 1))); do
if [ "$POLISH" = "1" ] && [ -f "$TMPDIR/polished_${i}.txt" ]; then cat "$TMPDIR/transcript_${i}.txt"
cat "$TMPDIR/polished_${i}.txt"
else
cat "$TMPDIR/transcript_${i}.txt"
fi
echo "" echo ""
done done
} > "$OUTPUT" } > "$OUTPUT"
+16 -16
View File
@@ -1,28 +1,28 @@
--- ---
name: agent-reach name: agent-reach
description: > description: >
MUST USE when user asks to search, browse, read, or interact with content from any of these platforms: Give your AI agent eyes to see the entire internet.
小红书/xiaohongshu/xhs, Twitter/推特/X, B站/bilibili, 17 platforms via CLI, MCP, curl, and Python scripts.
V2EX, Reddit, LinkedIn/领英, YouTube, GitHub code search, Zero config for 8 channels.
小宇宙播客, 雪球/股票行情, RSS feeds, or any web URL.
Also MUST USE for: web搜索/搜/查/找/look up/research, 招聘/求职/jobs, 分享的链接/URL.
Routes to CLI tools: xhs-cli, twitter-cli, rdt-cli, gh, yt-dlp, curl+Jina, mcporter.
13 platforms. Zero config for 6 channels.
【路由方式】SKILL.md 包含路由表和常用命令,复杂场景需按需阅读对应分类的 references/*.md。 【路由方式】SKILL.md 包含路由表和常用命令,复杂场景需按需阅读对应分类的 references/*.md。
分类:search / social (小红书/推特/B站/V2EX/Reddit) / career(LinkedIn) / dev(github) / web(网页/文章/RSS) / video(YouTube/B站/播客) 分类:search / social (小红书/抖音/微博/推特/B站/V2EX/Reddit) / career(LinkedIn) / dev(github) / web(网页/文章/公众号/RSS) / video(YouTube/B站/播客).
Use when user asks to search, read, or interact on any supported platform,
shares a URL, or asks to search the web.
triggers: triggers:
- search: 搜/查/找/search/搜索/查一下/帮我搜 - search: 搜/查/找/search/搜索/查一下/帮我搜
- social: - social:
- 小红书: xiaohongshu/xhs/小红书/红书 - 小红书: xiaohongshu/xhs/小红书/红书
- 抖音: douyin/抖音
- Twitter: twitter/推特/x.com/推文 - Twitter: twitter/推特/x.com/推文
- 微博: weibo/微博
- B站: bilibili/b站/哔哩哔哩 - B站: bilibili/b站/哔哩哔哩
- V2EX: v2ex - V2EX: v2ex
- Reddit: reddit - Reddit: reddit
- career: 招聘/职位/求职/linkedin/领英/找工作 - career: 招聘/职位/求职/linkedin/领英/找工作
- dev: github/代码/仓库/gh/issue/pr/分支/commit - dev: github/代码/仓库/gh/issue/pr/分支/commit
- web: 网页/链接/文章/rss/读一下/打开这个 - web: 网页/链接/文章/公众号/微信文章/rss/读一下/打开这个
- video: youtube/视频/播客/字幕/小宇宙/转录/yt - video: youtube/视频/播客/字幕/小宇宙/转录/yt
- finance: 雪球/股票/stock/xueqiu/行情/基金 - finance: 雪球/股票/stock/xueqiu/行情/基金
metadata: metadata:
@@ -32,17 +32,17 @@ metadata:
# Agent Reach — 路由器 # Agent Reach — 路由器
13 平台工具集合。根据用户意图选择对应分类。 17 平台工具集合。根据用户意图选择对应分类。
## 路由表 ## 路由表
| 用户意图 | 分类 | 详细文档 | | 用户意图 | 分类 | 详细文档 |
|---------|------|---------| |---------|------|---------|
| 网页搜索/代码搜索 | search | [references/search.md](references/search.md) | | 网页搜索/代码搜索 | search | [references/search.md](references/search.md) |
| 小红书/推特/B站/V2EX/Reddit | social | [references/social.md](references/social.md) | | 小红书/抖音/微博/推特/B站/V2EX/Reddit | social | [references/social.md](references/social.md) |
| 招聘/职位/LinkedIn | career | [references/career.md](references/career.md) | | 招聘/职位/LinkedIn | career | [references/career.md](references/career.md) |
| GitHub/代码 | dev | [references/dev.md](references/dev.md) | | GitHub/代码 | dev | [references/dev.md](references/dev.md) |
| 网页/文章/RSS | web | [references/web.md](references/web.md) | | 网页/文章/公众号/RSS | web | [references/web.md](references/web.md) |
| YouTube/B站/播客字幕 | video | [references/video.md](references/video.md) | | YouTube/B站/播客字幕 | video | [references/video.md](references/video.md) |
## 零配置快速命令 ## 零配置快速命令
@@ -58,7 +58,7 @@ curl -s "https://r.jina.ai/URL"
gh search repos "query" --sort stars --limit 10 gh search repos "query" --sort stars --limit 10
# Twitter 搜索 # Twitter 搜索
twitter search "query" -n 10 twitter search "query" --limit 10
# YouTube/B站字幕 # YouTube/B站字幕
yt-dlp --write-sub --skip-download -o "/tmp/%(id)s" "URL" yt-dlp --write-sub --skip-download -o "/tmp/%(id)s" "URL"
@@ -92,10 +92,10 @@ mcporter_list_servers()
根据用户需求,阅读对应的详细文档: 根据用户需求,阅读对应的详细文档:
- [搜索工具](references/search.md) — Exa AI 搜索 - [搜索工具](references/search.md) — Exa AI 搜索
- [社交媒体](references/social.md) — 小红书, Twitter, B站, V2EX, Reddit - [社交媒体](references/social.md) — 小红书, 抖音, Twitter, B站, V2EX, Reddit
- [职场招聘](references/career.md) — LinkedIn - [职场招聘](references/career.md) — LinkedIn
- [开发工具](references/dev.md) — GitHub CLI - [开发工具](references/dev.md) — GitHub CLI
- [网页阅读](references/web.md) — Jina Reader, RSS - [网页阅读](references/web.md) — Jina Reader, 微信公众号, RSS
- [视频播客](references/video.md) — YouTube, B站, 小宇宙 - [视频播客](references/video.md) — YouTube, B站, 小宇宙
## 配置渠道 ## 配置渠道
+85 -19
View File
@@ -1,18 +1,18 @@
--- ---
name: agent-reach name: agent-reach
description: > description: >
MUST USE when user asks to search, browse, read, or interact with content from any supported platform: Give your AI agent eyes to see the entire internet.
Twitter/X, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu, Search and read 17 platforms: Twitter/X, Reddit, YouTube, GitHub, Bilibili,
Xiaoyuzhou Podcast, LinkedIn, V2EX, Xueqiu (stocks), RSS, or any web URL. XiaoHongShu, Douyin, Weibo, WeChat Articles, Xiaoyuzhou Podcast, LinkedIn,
V2EX, Xueqiu, RSS, Exa web search, and any web page.
Also MUST USE for: web search, look up, research, find, share a URL/link, jobs/recruiting. Zero config for 8 channels. Use when the user asks to search, read, or interact
Routes to CLI tools: xhs-cli, twitter-cli, rdt-cli, gh, yt-dlp, curl+Jina, mcporter. on any supported platform, shares a URL, or asks to search the web.
13 platforms, zero config for 6 channels.
Triggers: "search twitter", "search xiaohongshu", "watch this video", Triggers: "search twitter", "search xiaohongshu", "watch this video",
"search the web", "look this up", "research", "youtube transcript", "search the web", "look this up", "research", "youtube transcript",
"search reddit", "read this link", "bilibili", "V2EX", "search reddit", "read this link", "bilibili", "douyin video",
"xiaoyuzhou", "podcast", "xueqiu", "stock quote", "雪球", "股票". "wechat article", "wechat official account", "weibo", "V2EX",
"xiaoyuzhou", "podcast", "xueqiu", "stock quote",
"install agent reach".
metadata: metadata:
openclaw: openclaw:
homepage: https://github.com/Panniantong/Agent-Reach homepage: https://github.com/Panniantong/Agent-Reach
@@ -20,7 +20,7 @@ metadata:
# Agent Reach — Usage Guide # Agent Reach — Usage Guide
Upstream tools for 13 platforms. Call them directly. Upstream tools for 17 platforms. Call them directly.
Run `agent-reach doctor` to check which channels are available. Run `agent-reach doctor` to check which channels are available.
@@ -41,18 +41,15 @@ mcporter call 'exa.web_search_exa(query: "query", numResults: 5)'
mcporter call 'exa.get_code_context_exa(query: "code question", tokensNum: 3000)' mcporter call 'exa.get_code_context_exa(query: "code question", tokensNum: 3000)'
``` ```
## Twitter/X (twitter-cli) ## Twitter/X (bird)
```bash ```bash
twitter -c search "query" -n 10 # search (-c = compact JSON, LLM-friendly) bird search "query" -n 10 # search
twitter -c tweet URL_OR_ID # read tweet + replies (supports /status/ URLs) bird read URL_OR_ID # read tweet (supports /status/ and /article/ URLs)
twitter -c article URL_OR_ID # read a Twitter Article bird user-tweets @username -n 20 # user timeline
twitter -c user-posts @username -n 20 # user timeline bird thread URL_OR_ID # full thread
twitter -c feed -n 20 # home timeline
``` ```
> Binary is `twitter` (`pipx install twitter-cli`, ≥ 0.8.5). The `bird` name in older docs has been retired. If `search` returns 404, run `pipx upgrade twitter-cli`.
## YouTube (yt-dlp) ## YouTube (yt-dlp)
```bash ```bash
@@ -108,6 +105,75 @@ mcporter call 'xiaohongshu.publish_content(title: "Title", content: "Body text",
> ``` > ```
> This keeps only: title, content, author, engagement counts, image URLs, and tags. > This keeps only: title, content, author, engagement counts, image URLs, and tags.
## Douyin (mcporter)
```bash
mcporter call 'douyin.parse_douyin_video_info(share_link: "https://v.douyin.com/xxx/")'
mcporter call 'douyin.get_douyin_download_link(share_link: "https://v.douyin.com/xxx/")'
```
> No login needed.
## WeChat Articles
**Search** (`miku_ai`):
```bash
# miku_ai is installed inside the agent-reach Python environment.
# Use the same interpreter that runs agent-reach (handles pipx / venv installs):
AGENT_REACH_PYTHON=$(python3 -c "import agent_reach, sys; print(sys.executable)" 2>/dev/null || echo python3)
$AGENT_REACH_PYTHON -c "
import asyncio
from miku_ai import get_wexin_article
async def s():
for a in await get_wexin_article('query', 5):
print(f'{a[\"title\"]} | {a[\"url\"]}')
asyncio.run(s())
"
```
**Read** (Camoufox — bypasses WeChat anti-bot):
```bash
cd ~/.agent-reach/tools/wechat-article-for-ai && python3 main.py "https://mp.weixin.qq.com/s/ARTICLE_ID"
```
> WeChat articles cannot be read with Jina Reader or curl. Use Camoufox.
## Weibo (mcporter)
```bash
# Trending topics
mcporter call 'weibo.get_trendings(limit: 20)'
# Search users
mcporter call 'weibo.search_users(keyword: "Lei Jun", limit: 10)'
# Get a user profile
mcporter call 'weibo.get_profile(uid: "1195230310")'
# Get a user's feed
mcporter call 'weibo.get_feeds(uid: "1195230310", limit: 20)'
# Get a user's hot posts
mcporter call 'weibo.get_hot_feeds(uid: "1195230310", limit: 10)'
# Search post content
mcporter call 'weibo.search_content(keyword: "artificial intelligence", limit: 20)'
# Search topics
mcporter call 'weibo.search_topics(keyword: "AI", limit: 10)'
# Get post comments
mcporter call 'weibo.get_comments(mid: "5099916367123456", limit: 50)'
# Get fans
mcporter call 'weibo.get_fans(uid: "1195230310", limit: 20)'
# Get followings
mcporter call 'weibo.get_followers(uid: "1195230310", limit: 20)'
```
> Zero config. No login needed. Uses the mobile API with auto-generated visitor cookies.
## Xiaoyuzhou Podcast (groq-whisper + ffmpeg) ## Xiaoyuzhou Podcast (groq-whisper + ffmpeg)
```bash ```bash
+70 -46
View File
@@ -1,67 +1,82 @@
# 社交媒体 & 社区 # 社交媒体 & 社区
小红书、Twitter/X、B站、V2EX、Reddit。 小红书、抖音、Twitter/X、微博、B站、V2EX、Reddit。
## 小红书 / XiaoHongShu(多后端) ## 小红书 / XiaoHongShu (xhs-cli)
小红书有三个后端,**先跑 `agent-reach doctor --json` 看 xiaohongshu 的 `active_backend` 是哪个**,再用对应命令组。 ### 稳定可用的命令
### 后端 A:OpenCLI(桌面首选,复用浏览器登录态)
```bash ```bash
# 搜索笔记 # 搜索笔记(推荐入口)
opencli xiaohongshu search "query" -f yaml xhs search "query"
# 读笔记正文+互动数据(用搜索结果里的完整 URL,含 xsec_token # 读笔记详情(必须用搜索结果中的 URL 或 ID,不能裸 note_id
opencli xiaohongshu note "NOTE_URL" -f yaml xhs read NOTE_ID_OR_URL
# 评论(支持楼中楼) # 查看评论
opencli xiaohongshu comments NOTE_ID -f yaml xhs comments NOTE_ID_OR_URL
# 首页推荐 feed # 浏览热门
opencli xiaohongshu feed -f yaml xhs hot
# 用户主页公开笔记 # 推荐 feed
opencli xiaohongshu user USER_ID -f yaml xhs feed
``` ```
> 要求 Chrome 打开且装了 OpenCLI 扩展。报 AUTH_REQUIRED 说明浏览器里没登录小红书,让用户在 Chrome 里登录一次即可。 ### 已知不稳定的命令(v0.6.4)
### 后端 Bxiaohongshu-mcp(服务器场景)
```bash ```bash
# 未登录时:先查状态,再取二维码给用户扫 # 以下命令当前可能返回 API error,谨慎使用:
mcporter call 'xiaohongshu.check_login_status()' --timeout 120000 xhs user USER_ID # 可能返回 {code: -1}
mcporter call 'xiaohongshu.get_login_qrcode()' --timeout 120000 xhs user-posts USER_ID # 可能返回 {code: -1}
xhs favorites # 可能返回 API error
# 搜索
mcporter call 'xiaohongshu.search_feeds(keyword: "query")' --timeout 120000
# 笔记详情+评论(feed_id 和 xsec_token 从搜索结果取)
mcporter call 'xiaohongshu.get_feed_detail(feed_id: "...", xsec_token: "...")' --timeout 120000
``` ```
> 首次调用会自动下载约 150MB 无头浏览器,务必带 `--timeout 120000`。未登录时 search 会挂死,先 check_login_status。 ### 重要注意事项
### 后端 Cxhs-cli(存量备选,上游 2026-03 起停更) > **安装**: `pipx install xiaohongshu-cli`,然后 `xhs login`(自动从浏览器提取 Cookie)。
```bash
xhs search "query" # 搜索
xhs read NOTE_ID_OR_URL # 读笔记(必须用搜索结果中的 URL/ID,不能裸 note_id
xhs comments NOTE_ID_OR_URL # 评论
xhs hot # 热门
xhs feed # 推荐
```
> 已知不稳定:`xhs user` / `xhs user-posts` / `xhs favorites` 可能返回 API error(上游停更无人修)。新装用户建议直接走后端 A/B。
### 通用注意事项
> **xsec_token 限制**: 小红书强制 xsec_token 机制,**不能直接用裸 note_id 去读**。正确流程:先搜索/feed 拿结果,再用结果中的完整 URL/ID 去读。三个后端都一样。
> >
> **频率控制**: 高频请求(批量搜索、深翻评论)会触发验证码,平台限制无法绕过。每次操作间隔 2-3 秒 > **xsec_token 限制**: 小红书强制 xsec_token 机制,**不能直接用裸 note_id 去读**。正确流程是:先 `xhs search``xhs feed` 获取结果,再用结果中的 URL/ID 去 `xhs read`。直接构造 note_id 会被拦截
> >
> **写操作(发帖/评论/点赞)**: 建议只读。xhs-cli v0.6.x 写操作可能因签名问题返回 406 > **频率控制**: 高频请求(批量搜索、深翻评论)会触发验证码,这是平台限制无法绕过。建议每次操作间隔 2-3 秒
>
> **POST 操作风险**: 发帖(post)、评论(comment)、点赞(like) 等写操作在 v0.6.x 可能因签名问题返回 406。如需使用,建议降级到 v0.3.5 (`pipx install xiaohongshu-cli==0.3.5`)。
## 抖音 / Douyin
### 安装与配置
`douyin-mcp-server` 是 **stdio 模式**的 MCP server,需先安装再注册到 mcporter:
```bash
# 1. 安装
pipx install douyin-mcp-server
# 2. 查找安装路径
pipx runpip douyin-mcp-server show -f 2>/dev/null | grep "Location" \
|| find ~/.local -name "douyin-mcp-server" 2>/dev/null | head -1
# 3. 注册到 mcporter(使用 stdio 模式,将路径替换为上一步的输出)
mcporter config add douyin --command "/path/to/douyin-mcp-server" --scope home
```
> **注意**`agent-reach install --channels douyin` 暂不支持抖音渠道(抖音在"可选渠道待解锁"列表)。
> HTTP 模式(`mcporter config add douyin http://localhost:18070/mcp`)**无法正常工作**,请使用上方 stdio 方式。
### 用法
```bash
# 解析视频信息
mcporter call 'douyin.parse_douyin_video_info(share_link: "https://v.douyin.com/xxx/")'
# 获取无水印下载链接
mcporter call 'douyin.get_douyin_download_link(share_link: "https://v.douyin.com/xxx/")'
# 提取视频文案
mcporter call 'douyin.extract_douyin_text(share_link: "https://v.douyin.com/xxx/")'
```
> **无需登录**
## Twitter/X (twitter-cli) ## Twitter/X (twitter-cli)
@@ -107,6 +122,15 @@ twitter likes
> >
> **输出格式**: 建议用 `--yaml``--json` 获得结构化输出,对 AI agent 更友好。 > **输出格式**: 建议用 `--yaml``--json` 获得结构化输出,对 AI agent 更友好。
## 微博 / Weibo
```bash
# 使用 Jina Reader 读取
curl -s "https://r.jina.ai/https://weibo.com/USER_ID/POST_ID"
```
> 微博主要通过网页抓取,推荐使用通用网页读取方式。
## B站 / Bilibili ## B站 / Bilibili
```bash ```bash
@@ -199,6 +223,6 @@ rdt popular --limit 10
rdt all --limit 10 rdt all --limit 10
``` ```
> **安装**: `pipx install 'git+https://github.com/public-clis/rdt-cli.git'`(PyPI 版本暂时落后,需从 GitHub 装 v0.4.2+)。需要先登录(`rdt login`)才能搜索和阅读。 > **安装**: `pipx install rdt-cli`(确保 v0.4.2+)。无需登录即可搜索和阅读。
> 需要登录的功能:`rdt feed --subs-only`(订阅列表)、`rdt saved`(收藏)。 > 需要登录的功能:`rdt feed --subs-only`(订阅列表)、`rdt saved`(收藏)。
> 建议使用 `--yaml` 输出,对 AI agent 更友好。 > 建议使用 `--yaml` 输出,对 AI agent 更友好。
+16 -17
View File
@@ -39,17 +39,6 @@ yt-dlp --dump-json "ytsearch5:query"
> **字幕注意**: 手动上传的字幕提取可靠;自动生成字幕可能存在行间重复,需后处理。 > **字幕注意**: 手动上传的字幕提取可靠;自动生成字幕可能存在行间重复,需后处理。
> **评论注意**: `--write-comments` 基于网页抓取(非 YouTube Data API),部分评论可能丢失。 > **评论注意**: `--write-comments` 基于网页抓取(非 YouTube Data API),部分评论可能丢失。
### 无字幕兜底:Whisper 音频转写
```bash
# 视频没有字幕时的兜底:下载音频并用 Whisper 转写(Groq 免费 key 即可)
agent-reach transcribe "https://www.youtube.com/watch?v=VIDEO_ID"
agent-reach transcribe ./local_audio.mp3 -o /tmp/transcript.txt
```
> 需要先配置 key`agent-reach configure groq-key gsk_xxx`(免费,console.groq.com
> 或 `agent-reach configure openai-key sk-xxx`。默认 auto 模式:groq 失败自动降级 openai。
## B站 / Bilibili (yt-dlp + bili-cli) ## B站 / Bilibili (yt-dlp + bili-cli)
### 视频元数据 (yt-dlp) ### 视频元数据 (yt-dlp)
@@ -82,15 +71,13 @@ bili rank -n 10
## 小宇宙播客 / Xiaoyuzhou Podcast ## 小宇宙播客 / Xiaoyuzhou Podcast
### 转录单集播客(可选 --polish 增强标点) ### 转录单集播客
```bash ```bash
# 输出 Markdown 文件到 /tmp/。--polish 让 Llama 3.3 70B 给文稿补中文标点+合理分段 # 输出 Markdown 文件到 /tmp/
~/.agent-reach/tools/xiaoyuzhou/transcribe.sh --polish "https://www.xiaoyuzhoufm.com/episode/EPISODE_ID" ~/.agent-reach/tools/xiaoyuzhou/transcribe.sh "https://www.xiaoyuzhoufm.com/episode/EPISODE_ID"
``` ```
> 转写 prompt 已要求 Whisper 输出中文标点;若标点效果仍不理想,可加 `--polish` 用 Groq 上免费的 Llama 3.3 70B 补标点+合理分段(9 分钟播客约多 ~7 秒)。每次转写多一轮 LLM 调用,按需使用。
### 前置要求 ### 前置要求
1. **ffmpeg**: `brew install ffmpeg` 1. **ffmpeg**: `brew install ffmpeg`
@@ -106,6 +93,18 @@ agent-reach doctor
> 输出 Markdown 文件默认保存到 `/tmp/` > 输出 Markdown 文件默认保存到 `/tmp/`
## 抖音视频解析
```bash
# 解析视频信息
mcporter call 'douyin.parse_douyin_video_info(share_link: "https://v.douyin.com/xxx/")'
# 获取无水印下载链接
mcporter call 'douyin.get_douyin_download_link(share_link: "https://v.douyin.com/xxx/")'
```
> 详见 [social.md](social.md#抖音--douyin)
## 选择指南 ## 选择指南
| 场景 | 推荐工具 | | 场景 | 推荐工具 |
@@ -113,4 +112,4 @@ agent-reach doctor
| YouTube 字幕 | yt-dlp | | YouTube 字幕 | yt-dlp |
| B站字幕 | yt-dlp | | B站字幕 | yt-dlp |
| 播客转录 | 小宇宙 transcribe.sh | | 播客转录 | 小宇宙 transcribe.sh |
| 无字幕音视频 | agent-reach transcribe | | 音视频解析 | douyin MCP |
+27 -1
View File
@@ -1,6 +1,6 @@
# 网页阅读 # 网页阅读
通用网页、RSS。 通用网页、微信公众号、RSS。
## 通用网页 (Jina Reader) ## 通用网页 (Jina Reader)
@@ -29,6 +29,30 @@ mcporter call 'web-reader.webReader(url: "https://example.com", return_format: "
**适用场景**: 需要更精确控制输出格式时使用。 **适用场景**: 需要更精确控制输出格式时使用。
## 微信公众号 / WeChat Articles
### 搜索公众号文章(通过 Exa
```bash
# 搜索微信公众号文章
mcporter call 'exa.web_search_exa(query: "搜索关键词", numResults: 5, includeDomains: ["mp.weixin.qq.com"])'
```
### 阅读公众号文章全文(通过 Exa)
```bash
# 抓取文章全文
mcporter call 'exa.crawling_exa(urls: ["https://mp.weixin.qq.com/s/ARTICLE_ID"], maxCharacters: 10000)'
```
### 可选:Camoufox 阅读(反爬更强)
```bash
cd ~/.agent-reach/tools/wechat-article-for-ai && python3 main.py "https://mp.weixin.qq.com/s/ARTICLE_ID"
```
> **注意**: Jina Reader 无法读取微信文章(被 CAPTCHA 拦截),推荐用 Exa。
## RSS (feedparser) ## RSS (feedparser)
```python ```python
@@ -47,4 +71,6 @@ for e in feedparser.parse('FEED_URL').entries[:5]:
|-----|---------| |-----|---------|
| 通用网页 | Jina Reader (`curl r.jina.ai`) | | 通用网页 | Jina Reader (`curl r.jina.ai`) |
| 需要图片/格式控制 | web-reader MCP | | 需要图片/格式控制 | web-reader MCP |
| 微信公众号 | Exa (搜索+阅读) / Camoufox (可选阅读) |
| RSS 订阅 | feedparser | | RSS 订阅 | feedparser |
| 微博/知乎等 | Jina Reader |
-253
View File
@@ -1,253 +0,0 @@
# -*- coding: utf-8 -*-
"""Whisper audio transcription with Groq → OpenAI fallback.
Downloads audio (yt-dlp), compresses + chunks (ffmpeg), and posts to a
Whisper-compatible API. Defaults to Groq's free `whisper-large-v3` and falls
back to OpenAI's `whisper-1` on HTTP error.
Public entry point:
transcribe(source, *, provider="auto", out_dir=None, config=None) -> str
Designed to be importable from channels (e.g. YouTubeChannel.transcribe).
"""
from __future__ import annotations
import shutil
import subprocess
import tempfile
from pathlib import Path
from typing import List, Optional
import requests
from agent_reach.config import Config
# Whisper API limit is 25MB; leave headroom for multipart overhead.
SIZE_LIMIT_BYTES = 24 * 1024 * 1024
CHUNK_SECONDS = 600 # 10 min — small enough that boundary cuts rarely lose meaning
PROVIDERS = {
"groq": {
"endpoint": "https://api.groq.com/openai/v1/audio/transcriptions",
"model": "whisper-large-v3",
"key_field": "groq_api_key",
},
"openai": {
"endpoint": "https://api.openai.com/v1/audio/transcriptions",
"model": "whisper-1",
"key_field": "openai_api_key",
},
}
class TranscribeError(RuntimeError):
"""Raised when transcription cannot complete."""
class MissingDependency(TranscribeError):
"""Raised when a required external binary is missing."""
class NoProviderConfigured(TranscribeError):
"""Raised when no provider has an API key configured."""
def _require(binary: str) -> None:
if not shutil.which(binary):
raise MissingDependency(f"{binary} not found in PATH")
def _run(cmd: List[str]) -> None:
"""Run a subprocess, raising TranscribeError on nonzero exit."""
proc = subprocess.run(cmd, capture_output=True, text=True)
if proc.returncode != 0:
raise TranscribeError(
f"{cmd[0]} failed (exit {proc.returncode}): {proc.stderr.strip()[:300]}"
)
def download_audio(url: str, out_dir: Path) -> Path:
"""Download audio with yt-dlp into out_dir; return the resulting file path."""
_require("yt-dlp")
template = out_dir / "source.%(ext)s"
_run(
[
"yt-dlp",
"-x",
"--audio-format",
"m4a",
"--audio-quality",
"0",
"-o",
str(template),
url,
]
)
files = sorted(out_dir.glob("source.*"))
if not files:
raise TranscribeError("yt-dlp produced no output file")
return files[0]
def compress_audio(src: Path, out_dir: Path) -> Path:
"""Re-encode to mono / 16kHz / 32kbps m4a — keeps most content under 25MB."""
_require("ffmpeg")
dst = out_dir / "compressed.m4a"
_run(
[
"ffmpeg",
"-loglevel",
"error",
"-y",
"-i",
str(src),
"-vn",
"-ac",
"1",
"-ar",
"16000",
"-b:a",
"32k",
str(dst),
]
)
return dst
def chunk_audio(src: Path, out_dir: Path, segment_seconds: int = CHUNK_SECONDS) -> List[Path]:
"""Split src into segments. Re-encodes each segment so cuts align to keyframes."""
_require("ffmpeg")
pattern = out_dir / "chunk_%03d.m4a"
_run(
[
"ffmpeg",
"-loglevel",
"error",
"-y",
"-i",
str(src),
"-f",
"segment",
"-segment_time",
str(segment_seconds),
"-ac",
"1",
"-ar",
"16000",
"-b:a",
"32k",
str(pattern),
]
)
chunks = sorted(out_dir.glob("chunk_*.m4a"))
if not chunks:
raise TranscribeError("ffmpeg produced no chunks")
return chunks
def _provider_key(provider: str, config: Config) -> Optional[str]:
field = PROVIDERS[provider]["key_field"]
val = config.get(field)
return val or None
def transcribe_chunk(
chunk: Path,
provider: str,
*,
config: Optional[Config] = None,
timeout: int = 120,
) -> str:
"""Transcribe one chunk via the named provider. Raises TranscribeError on failure."""
if provider not in PROVIDERS:
raise TranscribeError(f"unknown provider: {provider}")
cfg = config or Config()
key = _provider_key(provider, cfg)
if not key:
raise NoProviderConfigured(
f"{provider}: missing {PROVIDERS[provider]['key_field']} "
f"(configure with `agent-reach configure {provider}-key ...`)"
)
info = PROVIDERS[provider]
with chunk.open("rb") as fh:
try:
resp = requests.post(
info["endpoint"],
headers={"Authorization": f"Bearer {key}"},
files={"file": (chunk.name, fh, "audio/m4a")},
data={"model": info["model"], "response_format": "text"},
timeout=timeout,
)
except requests.RequestException as e:
raise TranscribeError(f"{provider}: network error: {e}") from e
if not resp.ok:
raise TranscribeError(f"{provider}: HTTP {resp.status_code}: {resp.text[:300]}")
return resp.text
def _provider_order(provider: str) -> List[str]:
if provider == "auto":
return ["groq", "openai"]
if provider in PROVIDERS:
return [provider]
raise TranscribeError(f"unknown provider: {provider} (use groq|openai|auto)")
def transcribe(
source: str,
*,
provider: str = "auto",
out_dir: Optional[Path] = None,
config: Optional[Config] = None,
) -> str:
"""Transcribe a URL or local file path. Returns the joined transcript text.
`provider` is one of `auto` (groq openai), `groq`, or `openai`.
`out_dir` defaults to a fresh temp directory; intermediate files stay there.
"""
cfg = config or Config()
order = _provider_order(provider)
# Validate at least one provider is configured before doing expensive work.
if not any(_provider_key(p, cfg) for p in order):
names = ", ".join(PROVIDERS[p]["key_field"] for p in order)
raise NoProviderConfigured(f"no provider key configured (need one of: {names})")
work_dir = Path(out_dir) if out_dir else Path(tempfile.mkdtemp(prefix="transcribe-"))
work_dir.mkdir(parents=True, exist_ok=True)
src_path = Path(source)
if src_path.is_file():
audio = src_path
else:
audio = download_audio(source, work_dir)
compressed = compress_audio(audio, work_dir)
if compressed.stat().st_size <= SIZE_LIMIT_BYTES:
chunks = [compressed]
else:
chunks = chunk_audio(compressed, work_dir)
pieces: List[str] = []
for chunk in chunks:
text = _transcribe_with_fallback(chunk, order, cfg)
pieces.append(text.strip())
return "\n".join(p for p in pieces if p)
def _transcribe_with_fallback(chunk: Path, order: List[str], config: Config) -> str:
"""Try each provider in order; return first success or raise the last error."""
last_err: Optional[Exception] = None
for p in order:
if not _provider_key(p, config):
# Skip silently — caller already validated at least one is configured.
continue
try:
return transcribe_chunk(chunk, p, config=config)
except TranscribeError as e:
last_err = e
continue
raise TranscribeError(f"all providers failed for {chunk.name}: {last_err}")
-26
View File
@@ -1,26 +0,0 @@
"""Subprocess helpers for consistent cross-platform text handling."""
from __future__ import annotations
import os
from collections.abc import Mapping
UTF8_ENV = {
"PYTHONUTF8": "1",
"PYTHONIOENCODING": "utf-8",
}
def utf8_subprocess_env(base: Mapping[str, str] | None = None) -> dict[str, str]:
"""Return an environment that forces Python child processes into UTF-8 mode."""
env = dict(base or os.environ)
env.update(UTF8_ENV)
return env
def mcporter_utf8_env_args() -> list[str]:
"""Return mcporter --env arguments for UTF-8 Python stdio servers."""
args = []
for key, value in UTF8_ENV.items():
args.extend(["--env", f"{key}={value}"])
return args
+45 -16
View File
@@ -64,7 +64,10 @@ Update Agent Reach: https://raw.githubusercontent.com/Panniantong/agent-reach/ma
| 🌐 **Web** | Read | Zero config | Any URL → clean Markdown ([Jina Reader](https://github.com/jina-ai/reader) ⭐9.8K) | | 🌐 **Web** | Read | Zero config | Any URL → clean Markdown ([Jina Reader](https://github.com/jina-ai/reader) ⭐9.8K) |
| 🐦 **Twitter/X** | Read · Search | Cookie | Cookie unlocks search, timeline, tweet reading, articles ([twitter-cli](https://github.com/public-clis/twitter-cli)) | | 🐦 **Twitter/X** | Read · Search | Cookie | Cookie unlocks search, timeline, tweet reading, articles ([twitter-cli](https://github.com/public-clis/twitter-cli)) |
| 📕 **XiaoHongShu** | Read · Search · **Post · Comment · Like** | Cookie | `pipx install xiaohongshu-cli` + `xhs login` ([xhs-cli](https://github.com/jackwener/xiaohongshu-cli)) | | 📕 **XiaoHongShu** | Read · Search · **Post · Comment · Like** | Cookie | `pipx install xiaohongshu-cli` + `xhs login` ([xhs-cli](https://github.com/jackwener/xiaohongshu-cli)) |
| 🎵 **Douyin** | Video parsing · Watermark-free download | mcporter | Via [douyin-mcp-server](https://github.com/yzfly/douyin-mcp-server), no login needed |
| 💼 **LinkedIn** | Jina Reader (public pages) | Full profiles, companies, job search | Tell your Agent "help me set up LinkedIn" | | 💼 **LinkedIn** | Jina Reader (public pages) | Full profiles, companies, job search | Tell your Agent "help me set up LinkedIn" |
| 💬 **WeChat Articles** | Search + Read | Zero config | Search + read WeChat Official Account articles via Exa (zero config) + optional [Camoufox](https://github.com/daijro/camoufox) |
| 📰 **Weibo** | Trending · Search · Feeds · Comments | Zero config | Hot search, content/user/topic search, feeds, comments ([mcp-server-weibo](https://github.com/Panniantong/mcp-server-weibo)) |
| 💻 **V2EX** | Hot topics · Node topics · Topic detail + replies · User profile | Zero config | Public JSON API, no auth required. Great for tech community content | | 💻 **V2EX** | Hot topics · Node topics · Topic detail + replies · User profile | Zero config | Public JSON API, no auth required. Great for tech community content |
| 📈 **Xueqiu (雪球)** | Stock quotes · Search · Hot posts · Hot stocks | Browser cookie | Tell your Agent "help me set up Xueqiu" | | 📈 **Xueqiu (雪球)** | Stock quotes · Search · Hot posts · Hot stocks | Browser cookie | Tell your Agent "help me set up Xueqiu" |
| 🎙️ **Xiaoyuzhou Podcast** | Transcription | Free API key | Podcast audio → full text transcript via Groq Whisper (free) | | 🎙️ **Xiaoyuzhou Podcast** | Transcription | Free API key | Podcast audio → full text transcript via Groq Whisper (free) |
@@ -81,15 +84,6 @@ Update Agent Reach: https://raw.githubusercontent.com/Panniantong/agent-reach/ma
## Quick Start ## Quick Start
> ⚠️ **OpenClaw users: enable `exec` permission first**
>
> Agent Reach relies on the Agent running shell commands (`pip install`, `mcporter`, `twitter`, etc.). If your OpenClaw uses the default `messaging` tool profile, the Agent won't be able to run them. **Enable `exec` before installing:**
>
> ```bash
> openclaw config set tools.profile "coding"
> ```
> Or set `"tools": { "profile": "coding" }` in `~/.openclaw/openclaw.json`. After changing it, restart the Gateway (`openclaw gateway restart`) and start a new conversation. Other platforms (Claude Code, Cursor, Windsurf, etc.) are not affected.
Copy this to your AI Agent (Claude Code, OpenClaw, Cursor, etc.): Copy this to your AI Agent (Claude Code, OpenClaw, Cursor, etc.):
``` ```
@@ -103,12 +97,6 @@ The Agent auto-installs, detects your environment, and tells you what's ready.
> Update Agent Reach: https://raw.githubusercontent.com/Panniantong/agent-reach/main/docs/update.md > Update Agent Reach: https://raw.githubusercontent.com/Panniantong/agent-reach/main/docs/update.md
> ``` > ```
> 🛡️ **Worried about security?** Use safe mode — it won't auto-install system packages, it only tells you what you need:
> ```
> Install Agent Reach (safe mode): https://raw.githubusercontent.com/Panniantong/agent-reach/main/docs/install.md
> Use the --safe flag during install
> ```
<details> <details>
<summary>Manual install</summary> <summary>Manual install</summary>
@@ -219,6 +207,7 @@ channels/
├── bilibili.py → yt-dlp ← swap to bilibili-api… ├── bilibili.py → yt-dlp ← swap to bilibili-api…
├── reddit.py → rdt-cli ← search + read, cookie auth required ├── reddit.py → rdt-cli ← search + read, cookie auth required
├── xiaohongshu.py → mcporter MCP ← swap to other XHS tools… ├── xiaohongshu.py → mcporter MCP ← swap to other XHS tools…
├── douyin.py → mcporter MCP ← swap to other Douyin tools…
├── linkedin.py → linkedin-mcp ← swap to LinkedIn API… ├── linkedin.py → linkedin-mcp ← swap to LinkedIn API…
├── rss.py → feedparser ← swap to atoma… ├── rss.py → feedparser ← swap to atoma…
├── exa_search.py → mcporter MCP ← swap to Tavily, SerpAPI… ├── exa_search.py → mcporter MCP ← swap to Tavily, SerpAPI…
@@ -240,7 +229,10 @@ Each channel file only checks whether its upstream tool is installed and working
| GitHub | [gh CLI](https://cli.github.com) | Official tool, full API after auth | | GitHub | [gh CLI](https://cli.github.com) | Official tool, full API after auth |
| Read RSS | [feedparser](https://github.com/kurtmckee/feedparser) | Python ecosystem standard, 2.3K stars | | Read RSS | [feedparser](https://github.com/kurtmckee/feedparser) | Python ecosystem standard, 2.3K stars |
| XiaoHongShu | [xhs-cli](https://github.com/jackwener/xiaohongshu-cli) | 1.5K stars, pipx install, search/read/comment/post | | XiaoHongShu | [xhs-cli](https://github.com/jackwener/xiaohongshu-cli) | 1.5K stars, pipx install, search/read/comment/post |
| Douyin | [douyin-mcp-server](https://github.com/yzfly/douyin-mcp-server) | MCP server, no login needed, video parsing + watermark-free download |
| LinkedIn | [linkedin-scraper-mcp](https://github.com/stickerdaniel/linkedin-mcp-server) | 1.2K stars, MCP server, browser automation | | LinkedIn | [linkedin-scraper-mcp](https://github.com/stickerdaniel/linkedin-mcp-server) | 1.2K stars, MCP server, browser automation |
| WeChat Articles | [Exa](https://exa.ai) (search + read) + [Camoufox](https://github.com/daijro/camoufox) (optional) | Zero-config search + full article reading |
| Weibo | `mcporter` | `mcporter call 'weibo.get_trendings(limit: 10)'` |
| Xiaoyuzhou Podcast | `transcribe.sh` | `bash ~/.agent-reach/tools/xiaoyuzhou/transcribe.sh <URL>` | | Xiaoyuzhou Podcast | `transcribe.sh` | `bash ~/.agent-reach/tools/xiaoyuzhou/transcribe.sh <URL>` |
> 📌 These are the *current* choices. Don't like one? Swap out the file. That's the whole point of scaffolding. > 📌 These are the *current* choices. Don't like one? Swap out the file. That's the whole point of scaffolding.
@@ -303,11 +295,44 @@ Agent Reach uses twitter-cli which accesses Twitter via cookie auth — same as
Install `pipx install xiaohongshu-cli`, then `xhs login` (auto-extracts cookies from browser). Your agent can then use `xhs search "query"` to search notes, `xhs read NOTE_ID` to read details, `xhs comments NOTE_ID` to view comments. No Docker needed. Install `pipx install xiaohongshu-cli`, then `xhs login` (auto-extracts cookies from browser). Your agent can then use `xhs search "query"` to search notes, `xhs read NOTE_ID` to read details, `xhs comments NOTE_ID` to view comments. No Docker needed.
</details> </details>
<details>
<summary><strong>How to parse Douyin / 抖音 videos with AI agent?</strong></summary>
Install douyin-mcp-server, then your agent can use `mcporter call 'douyin.parse_douyin_video_info(share_link: "share_url")'` to parse video info and get watermark-free download links. No login required — just share the Douyin link. See https://github.com/yzfly/douyin-mcp-server
</details>
<details>
<summary><strong>How to extract scripts from both Douyin and XiaoHongShu with one MCP?</strong></summary>
If you want one MCP server that can handle:
- Douyin videos
- XiaoHongShu video notes
- XiaoHongShu image notes
and directly write `script.md` + `info.json`, you can point the existing `douyin` mcporter alias at:
- https://github.com/JNHFlow21/social-post-extractor-mcp
It keeps backward compatibility with:
- `parse_douyin_video_info`
- `get_douyin_download_link`
- `extract_douyin_text`
and adds unified tools:
- `parse_social_post_info`
- `extract_social_post_script`
This is useful when your agent workflow is “paste a link, get a script file”.
</details>
--- ---
## Credits ## Credits
[twitter-cli](https://github.com/public-clis/twitter-cli) · [rdt-cli](https://github.com/public-clis/rdt-cli) · [xhs-cli](https://github.com/jackwener/xiaohongshu-cli) · [bili-cli](https://github.com/public-clis/bilibili-cli) · [yt-dlp](https://github.com/yt-dlp/yt-dlp) · [Jina Reader](https://github.com/jina-ai/reader) · [Exa](https://exa.ai) · [mcporter](https://github.com/nicobailon/mcporter) · [feedparser](https://github.com/kurtmckee/feedparser) · [linkedin-scraper-mcp](https://github.com/stickerdaniel/linkedin-mcp-server) [twitter-cli](https://github.com/public-clis/twitter-cli) · [rdt-cli](https://github.com/public-clis/rdt-cli) · [xhs-cli](https://github.com/jackwener/xiaohongshu-cli) · [bili-cli](https://github.com/public-clis/bilibili-cli) · [yt-dlp](https://github.com/yt-dlp/yt-dlp) · [Jina Reader](https://github.com/jina-ai/reader) · [Exa](https://exa.ai) · [mcporter](https://github.com/nicobailon/mcporter) · [feedparser](https://github.com/kurtmckee/feedparser) · [douyin-mcp-server](https://github.com/yzfly/douyin-mcp-server) · [linkedin-scraper-mcp](https://github.com/stickerdaniel/linkedin-mcp-server)
## Contact ## Contact
@@ -328,6 +353,10 @@ For collaboration or questions, add me on WeChat — I'll invite you to the comm
## Friends ## Friends
[FluxNode](https://fluxnode.org) — Low-cost AI API gateway, 90% off official pricing, pay-as-you-go or subscription. Works with OpenClaw, Claude Code, and any Agent.
[OpenClaw for Enterprise](https://github.com/littleben/openclaw-for-enterprise) — Enterprise-grade multi-user OpenClaw deployment, use AI directly in Feishu/Lark, container isolation, one-command management.
[OpenClaw on Tencent Cloud](https://www.tencentcloud.com/act/pro/intl-openclaw?referral_code=G76Y819A&lang=en&pg=) — One-click OpenClaw on Tencent Cloud: chat to connect Agent Reach & unlock internet power. [OpenClaw on Tencent Cloud](https://www.tencentcloud.com/act/pro/intl-openclaw?referral_code=G76Y819A&lang=en&pg=) — One-click OpenClaw on Tencent Cloud: chat to connect Agent Reach & unlock internet power.
## Star History ## Star History
+4
View File
@@ -348,6 +348,10 @@ douyin-mcp-server를 설치한 다음, 에이전트가 `mcporter call 'douyin.pa
## 관련 프로젝트 ## 관련 프로젝트
[FluxNode](https://fluxnode.org) — 저비용 AI API 게이트웨이, 공식 가격의 90% 할인, 종량제 또는 구독. OpenClaw, Claude Code 및 모든 에이전트와 호환.
[OpenClaw for Enterprise](https://github.com/littleben/openclaw-for-enterprise) — 엔터프라이즈급 다중 사용자 OpenClaw 배포, Feishu/Lark에서 AI 직접 사용, 컨테이너 격리, 원 명령어 관리.
[OpenClaw on Tencent Cloud](https://www.tencentcloud.com/act/pro/intl-openclaw?referral_code=G76Y819A&lang=en&pg=) — Tencent Cloud에서 원클릭 OpenClaw: 채팅으로 Agent Reach를 연결하고 인터넷 기능을 활성화하세요. [OpenClaw on Tencent Cloud](https://www.tencentcloud.com/act/pro/intl-openclaw?referral_code=G76Y819A&lang=en&pg=) — Tencent Cloud에서 원클릭 OpenClaw: 채팅으로 Agent Reach를 연결하고 인터넷 기능을 활성화하세요.
## Star History ## Star History
+91 -18
View File
@@ -40,7 +40,7 @@ All Agent Reach files go in dedicated directories — **never in the agent works
| Purpose | Directory | Example | | Purpose | Directory | Example |
|---------|-----------|---------| |---------|-----------|---------|
| Config & tokens | `~/.agent-reach/` | `~/.agent-reach/config.json` | | Config & tokens | `~/.agent-reach/` | `~/.agent-reach/config.json` |
| Upstream tool repos | `~/.agent-reach/tools/` | `~/.agent-reach/tools/xiaoyuzhou/` | | Upstream tool repos | `~/.agent-reach/tools/` | `~/.agent-reach/tools/douyin-mcp-server/` |
| Temporary files | `/tmp/` | `/tmp/yt-dlp-output/` | | Temporary files | `/tmp/` | `/tmp/yt-dlp-output/` |
| Skills | `~/.openclaw/skills/agent-reach/` | SKILL.md | | Skills | `~/.openclaw/skills/agent-reach/` | SKILL.md |
@@ -61,19 +61,6 @@ pip install https://github.com/Panniantong/agent-reach/archive/main.zip
agent-reach install --env=auto agent-reach install --env=auto
``` ```
> 💡 **Windows / Microsoft Store Python alias?**
> 如果 `python3 --version` 打开 Microsoft Store,或 `where python3` 指向
> `...\AppData\Local\Microsoft\WindowsApps\python3.exe`,说明 `python3` 是 Windows
> 的 Store alias,不是可用的 Python 安装。请改用 Python Launcher `py -3`,或实际安装目录里的 `python.exe`
>
> PowerShell 示例:
> ```powershell
> py -3 -m venv $env:USERPROFILE\.agent-reach-venv
> $env:USERPROFILE\.agent-reach-venv\Scripts\Activate.ps1
> python -m pip install https://github.com/Panniantong/agent-reach/archive/main.zip
> agent-reach install --env=auto
> ```
This installs core infrastructure (gh CLI, Node.js, mcporter, Exa search, yt-dlp config) and activates these zero-config channels: This installs core infrastructure (gh CLI, Node.js, mcporter, Exa search, yt-dlp config) and activates these zero-config channels:
- Web (Jina Reader), YouTube, GitHub, RSS, Exa Search, V2EX, Bilibili (basic) - Web (Jina Reader), YouTube, GitHub, RSS, Exa Search, V2EX, Bilibili (basic)
@@ -97,23 +84,26 @@ After installing the basics, **ask the user** which additional channels they nee
> 还有这些可选渠道,你需要哪些? > 还有这些可选渠道,你需要哪些?
> >
> - 🐦 **Twitter/X** — 搜推文、看时间线(需要登录 Cookie) > - 🐦 **Twitter/X** — 搜推文、看时间线(需要登录 Cookie)
> - 📰 **微博** — 热搜、搜索、用户动态
> - 💬 **微信公众号** — 搜索和阅读公众号文章
> - 📈 **雪球** — 股票行情、热门帖子(需要登录 Cookie) > - 📈 **雪球** — 股票行情、热门帖子(需要登录 Cookie)
> - 🎙️ **小宇宙播客** — 音频转文字(需要免费 Groq Key) > - 🎙️ **小宇宙播客** — 音频转文字(需要免费 Groq Key)
> - 📕 **小红书** — 阅读、搜索、发帖(需要登录) > - 📕 **小红书** — 阅读、搜索、发帖(需要登录)
> - 📖 **Reddit** — 搜索和阅读帖子 > - 📖 **Reddit** — 搜索和阅读帖子
> - 📺 **B站完整版** — 热门、排行、搜索 > - 📺 **B站完整版** — 热门、排行、搜索
> - 🎵 **抖音** — 视频解析
> - 💼 **LinkedIn** — Profile、职位搜索 > - 💼 **LinkedIn** — Profile、职位搜索
> >
> 告诉我你要哪些,比如"帮我装小红书和 Twitter"。或者说"全部装"。 > 告诉我你要哪些,比如"帮我装微博和 Twitter"。或者说"全部装"。
Based on the user's choice, run: Based on the user's choice, run:
```bash ```bash
agent-reach install --env=auto --channels=twitter,xiaohongshu # Example: user chose Twitter + XHS agent-reach install --env=auto --channels=twitter,weibo # Example: user chose Twitter + Weibo
agent-reach install --env=auto --channels=all # User wants everything agent-reach install --env=auto --channels=all # User wants everything
``` ```
Supported channel names: `twitter`, `xiaoyuzhou`, `xueqiu`, `xiaohongshu`, `reddit`, `bilibili`, `linkedin`, `all` Supported channel names: `twitter`, `weibo`, `wechat`, `xiaoyuzhou`, `xueqiu`, `xiaohongshu`, `reddit`, `bilibili`, `douyin`, `linkedin`, `all`
### Step 3: Fix what's broken ### Step 3: Fix what's broken
@@ -199,6 +189,18 @@ xhs login
> mcporter config add xiaohongshu http://localhost:18060/mcp > mcporter config add xiaohongshu http://localhost:18060/mcp
> ``` > ```
**微博 / Weibo (mcp-server-weibo):**
> "微博已默认安装,装好即用。可搜索微博内容、查看热搜、获取用户动态和评论。"
如果自动安装失败,手动安装:
```bash
pip install git+https://github.com/Panniantong/mcp-server-weibo.git
mcporter config add weibo --command 'mcp-server-weibo'
```
> 无需登录、无需 Cookie、无需代理。海外服务器也可以直接访问。
**雪球 / Xueqiu (股票行情 + 热门帖子):** **雪球 / Xueqiu (股票行情 + 热门帖子):**
> "雪球需要登录后的 Cookie。请先在 Chrome 里登录 xueqiu.com,然后运行:" > "雪球需要登录后的 Cookie。请先在 Chrome 里登录 xueqiu.com,然后运行:"
@@ -237,6 +239,75 @@ agent-reach configure groq-key gsk_xxxxx
> - 转录质量高(Whisper large-v3),但不区分说话人 > - 转录质量高(Whisper large-v3),但不区分说话人
> - 2 小时以上的播客建议分批处理 > - 2 小时以上的播客建议分批处理
**抖音 / Douyin (douyin-mcp-server):**
> "抖音视频解析需要一个 MCP 服务。安装 douyin-mcp-server 后即可解析视频、获取无水印下载链接。"
```bash
# 1. 安装
pip install douyin-mcp-server
# 2. 启动 HTTP 服务(端口 18070
# 方式一:用 uv(推荐)
mkdir -p ~/.agent-reach/tools && cd ~/.agent-reach/tools
git clone https://github.com/yzfly/douyin-mcp-server.git && cd douyin-mcp-server
uv sync && uv run python run_http.py
# 方式二:直接用 Python 启动
python -c "
from douyin_mcp_server.server import mcp
mcp.settings.host = '127.0.0.1'
mcp.settings.port = 18070
mcp.run(transport='streamable-http')
"
# 3. 注册到 mcporter
mcporter config add douyin http://localhost:18070/mcp
```
> 无需认证即可解析视频信息和获取下载链接。
> 如需 AI 语音识别提取文案功能,需要配置硅基流动 API Key(`export API_KEY="sk-xxx"`)。
>
> 详见 https://github.com/yzfly/douyin-mcp-server
**可选实现:Douyin + XiaoHongShu unified extractor**
> "如果你想把抖音和小红书统一成一个 MCP,并直接输出 `script.md``info.json`,可以改用 social-post-extractor-mcp。"
适用场景:
- 抖音视频转文字稿
- 小红书视频笔记转文字稿
- 小红书图文笔记正文 + 图片文字提取
兼容性:
- 仍然可以注册成 `douyin` 这个 mcporter server 名称
- 兼容旧工具名 `parse_douyin_video_info` / `get_douyin_download_link` / `extract_douyin_text`
- 同时新增 `parse_social_post_info` / `extract_social_post_script`
示例配置:
```bash
git clone https://github.com/JNHFlow21/social-post-extractor-mcp.git
cd social-post-extractor-mcp
uv sync
mcporter config add douyin \
--command /bin/zsh \
--arg -lc \
--arg "cd '$PWD' && exec '.venv/bin/python' -m social_post_extractor_mcp" \
--env ASR_PROVIDER=bailian \
--env ASR_MODEL=paraformer-v2 \
--env VISION_PROVIDER=bailian \
--env VISION_MODEL=qwen3-vl-flash \
--env CLEAN_PROVIDER=bailian \
--env CLEAN_MODEL=qwen-flash \
--env BAILIAN_API_KEY=YOUR_BAILIAN_API_KEY
```
> 这个实现更适合“把链接直接交给 Agent,然后拿到脚本文件”的工作流。
>
> 详见 https://github.com/JNHFlow21/social-post-extractor-mcp
**LinkedIn (可选 — linkedin-scraper-mcp):** **LinkedIn (可选 — linkedin-scraper-mcp):**
> "LinkedIn 基本内容可通过 Jina Reader 读取。完整功能(Profile 详情、职位搜索)需要 linkedin-scraper-mcp。" > "LinkedIn 基本内容可通过 Jina Reader 读取。完整功能(Profile 详情、职位搜索)需要 linkedin-scraper-mcp。"
@@ -304,7 +375,7 @@ If the user wants a different agent to handle it, let them choose.
| Command | What it does | | Command | What it does |
|---------|-------------| |---------|-------------|
| `agent-reach install --env=auto` | Install core channels (lightweight, zero-config) | | `agent-reach install --env=auto` | Install core channels (lightweight, zero-config) |
| `agent-reach install --env=auto --channels=twitter,xiaohongshu` | Install core + optional channels | | `agent-reach install --env=auto --channels=twitter,weibo` | Install core + optional channels |
| `agent-reach install --env=auto --channels=all` | Install everything | | `agent-reach install --env=auto --channels=all` | Install everything |
| `agent-reach install --env=auto --safe` | Safe setup (no auto system changes) | | `agent-reach install --env=auto --safe` | Safe setup (no auto system changes) |
| `agent-reach install --env=auto --dry-run` | Preview what would be done | | `agent-reach install --env=auto --dry-run` | Preview what would be done |
@@ -327,6 +398,8 @@ After installation, use upstream tools directly. See SKILL.md for the full comma
| Web | `curl` + Jina | `curl -s "https://r.jina.ai/URL"` | | Web | `curl` + Jina | `curl -s "https://r.jina.ai/URL"` |
| Exa Search | `mcporter` | `mcporter call 'exa.web_search_exa(...)'` | | Exa Search | `mcporter` | `mcporter call 'exa.web_search_exa(...)'` |
| 小红书 | `mcporter` | `mcporter call 'xiaohongshu.search_feeds(...)'` | | 小红书 | `mcporter` | `mcporter call 'xiaohongshu.search_feeds(...)'` |
| 微博 | `mcporter` | `mcporter call 'weibo.get_trendings(limit: 10)'` |
| 小宇宙播客 | `transcribe.sh` | `bash ~/.agent-reach/tools/xiaoyuzhou/transcribe.sh <URL>` | | 小宇宙播客 | `transcribe.sh` | `bash ~/.agent-reach/tools/xiaoyuzhou/transcribe.sh <URL>` |
| 抖音 | `mcporter` | `mcporter call 'douyin.parse_douyin_video_info(...)'` |
| LinkedIn | `mcporter` | `mcporter call 'linkedin.get_person_profile(...)'` | | LinkedIn | `mcporter` | `mcporter call 'linkedin.get_person_profile(...)'` |
| RSS | `feedparser` | `python3 -c "import feedparser; ..."` | | RSS | `feedparser` | `python3 -c "import feedparser; ..."` |
+2 -2
View File
@@ -50,8 +50,8 @@ Run these commands to ensure all upstream CLI tools are installed. Skip any that
# Twitter/X — twitter-cli (replaces deprecated bird CLI) # Twitter/X — twitter-cli (replaces deprecated bird CLI)
which twitter >/dev/null 2>&1 || pipx install twitter-cli 2>/dev/null || uv tool install twitter-cli 2>/dev/null which twitter >/dev/null 2>&1 || pipx install twitter-cli 2>/dev/null || uv tool install twitter-cli 2>/dev/null
# Reddit — rdt-cli (replaces Exa-based approach; PyPI lags, install from GitHub) # Reddit — rdt-cli (replaces Exa-based approach)
which rdt >/dev/null 2>&1 || pipx install 'git+https://github.com/public-clis/rdt-cli.git' 2>/dev/null || uv tool install --from 'git+https://github.com/public-clis/rdt-cli.git' rdt-cli 2>/dev/null which rdt >/dev/null 2>&1 || pipx install rdt-cli 2>/dev/null || uv tool install rdt-cli 2>/dev/null
# XiaoHongShu — xhs-cli (replaces Docker MCP) # XiaoHongShu — xhs-cli (replaces Docker MCP)
which xhs >/dev/null 2>&1 || pipx install xiaohongshu-cli 2>/dev/null || uv tool install xiaohongshu-cli 2>/dev/null which xhs >/dev/null 2>&1 || pipx install xiaohongshu-cli 2>/dev/null || uv tool install xiaohongshu-cli 2>/dev/null
+7 -1
View File
@@ -1,6 +1,6 @@
[project] [project]
name = "agent-reach" name = "agent-reach"
version = "1.4.2" version = "1.4.0"
description = "Give your AI Agent eyes to see the entire internet. Search + Read 10+ platforms." description = "Give your AI Agent eyes to see the entire internet. Search + Read 10+ platforms."
readme = "README.md" readme = "README.md"
license = {text = "MIT"} license = {text = "MIT"}
@@ -64,6 +64,12 @@ build-backend = "hatchling.build"
[tool.hatch.build.targets.wheel] [tool.hatch.build.targets.wheel]
packages = ["agent_reach"] packages = ["agent_reach"]
[tool.hatch.build.targets.wheel.force-include]
"agent_reach/guides" = "agent_reach/guides"
# Keep the whole skill directory so SKILL.md, SKILL_en.md, and references/ ship together.
"agent_reach/skill" = "agent_reach/skill"
"agent_reach/scripts" = "agent_reach/scripts"
[tool.ruff] [tool.ruff]
target-version = "py310" target-version = "py310"
line-length = 100 line-length = 100
+45 -73
View File
@@ -1,17 +1,10 @@
# -*- coding: utf-8 -*- # -*- coding: utf-8 -*-
"""Contract tests for channel adapters.""" """Contract tests for channel adapters."""
import subprocess
from agent_reach.channels import get_all_channels from agent_reach.channels import get_all_channels
from agent_reach.config import Config from agent_reach.config import Config
def _fake_run_ok(cmd, **kwargs):
"""Pretend any probed CLI executes fine and prints a version."""
return subprocess.CompletedProcess(cmd, 0, "2026.06.09", "")
def test_channel_registry_contract(): def test_channel_registry_contract():
channels = get_all_channels() channels = get_all_channels()
assert channels, "channel registry must not be empty" assert channels, "channel registry must not be empty"
@@ -36,67 +29,6 @@ def test_channel_check_contract_with_minimal_runtime(monkeypatch, tmp_path):
assert isinstance(message, str) and message.strip() assert isinstance(message, str) and message.strip()
def test_channel_active_backend_attribute_contract():
"""Every channel exposes active_backend (default None, or str once set)."""
for ch in get_all_channels():
assert hasattr(ch, "active_backend")
# Fresh instances must default to None / str (class attribute on base)
fresh = type(ch)()
assert fresh.active_backend is None or isinstance(fresh.active_backend, str)
def test_channel_active_backend_set_by_check(monkeypatch, tmp_path):
"""After check(), active_backend is None or a str — never anything else."""
monkeypatch.setattr("shutil.which", lambda _cmd: None)
# Keep the network-based channels (V2EX/Xueqiu/Bilibili API) deterministic.
import urllib.request
from urllib.error import URLError
def _no_net(*_a, **_k):
raise URLError("offline")
monkeypatch.setattr(urllib.request, "urlopen", _no_net)
import agent_reach.channels.xueqiu as xueqiu_mod
monkeypatch.setattr(xueqiu_mod, "_cookies_initialized", True)
monkeypatch.setattr(xueqiu_mod._opener, "open", _no_net)
config = Config(config_path=tmp_path / "config.yaml")
for ch in get_all_channels():
ch.check(config)
assert ch.active_backend is None or isinstance(ch.active_backend, str), (
f"{ch.name}: active_backend must be None or str after check()"
)
def test_ordered_backends_contract(tmp_path):
"""ordered_backends(config) is a reordering (same multiset) of backends."""
config = Config(config_path=tmp_path / "config.yaml")
for ch in get_all_channels():
ordered = ch.ordered_backends(config)
assert isinstance(ordered, list)
assert sorted(ordered) == sorted(ch.backends), (
f"{ch.name}: ordered_backends must be a permutation of backends"
)
# And without any config at all
ordered_none = ch.ordered_backends(None)
assert sorted(ordered_none) == sorted(ch.backends)
def test_ordered_backends_override_moves_backend_to_front():
"""Config key <channel>_backend promotes the named backend to front."""
from agent_reach.channels.twitter import TwitterChannel
ch = TwitterChannel()
ordered = ch.ordered_backends({"twitter_backend": "bird"})
assert ordered[0] == "bird CLI (legacy)"
assert sorted(ordered) == sorted(ch.backends)
# Unknown override is ignored — never hides working backends
ordered_unknown = ch.ordered_backends({"twitter_backend": "no-such-tool"})
assert ordered_unknown == list(ch.backends)
def test_youtube_warns_when_node_only_and_no_config(monkeypatch, tmp_path): def test_youtube_warns_when_node_only_and_no_config(monkeypatch, tmp_path):
"""YouTube should warn when only Node.js is installed but no yt-dlp config exists.""" """YouTube should warn when only Node.js is installed but no yt-dlp config exists."""
from agent_reach.channels.youtube import YouTubeChannel from agent_reach.channels.youtube import YouTubeChannel
@@ -109,7 +41,6 @@ def test_youtube_warns_when_node_only_and_no_config(monkeypatch, tmp_path):
return None # deno not installed return None # deno not installed
monkeypatch.setattr("shutil.which", fake_which) monkeypatch.setattr("shutil.which", fake_which)
monkeypatch.setattr("subprocess.run", _fake_run_ok) # yt-dlp probe really executes now
# Point to a non-existent config file # Point to a non-existent config file
monkeypatch.setattr("os.path.expanduser", lambda p: str(tmp_path / ".config/yt-dlp/config")) monkeypatch.setattr("os.path.expanduser", lambda p: str(tmp_path / ".config/yt-dlp/config"))
@@ -117,7 +48,6 @@ def test_youtube_warns_when_node_only_and_no_config(monkeypatch, tmp_path):
status, message = ch.check() status, message = ch.check()
assert status == "warn" assert status == "warn"
assert "--js-runtimes" in message assert "--js-runtimes" in message
assert ch.active_backend == "yt-dlp" # 本体活着,warn 只关乎 JS runtime
def test_youtube_warns_with_windows_specific_fix_command(monkeypatch, tmp_path): def test_youtube_warns_with_windows_specific_fix_command(monkeypatch, tmp_path):
@@ -132,7 +62,6 @@ def test_youtube_warns_with_windows_specific_fix_command(monkeypatch, tmp_path):
return None return None
monkeypatch.setattr("shutil.which", fake_which) monkeypatch.setattr("shutil.which", fake_which)
monkeypatch.setattr("subprocess.run", _fake_run_ok) # yt-dlp probe really executes now
monkeypatch.setattr("agent_reach.utils.paths.sys.platform", "win32") monkeypatch.setattr("agent_reach.utils.paths.sys.platform", "win32")
monkeypatch.setenv("APPDATA", str(tmp_path / "AppData" / "Roaming")) monkeypatch.setenv("APPDATA", str(tmp_path / "AppData" / "Roaming"))
@@ -155,12 +84,53 @@ def test_youtube_ok_when_deno_installed(monkeypatch):
return None return None
monkeypatch.setattr("shutil.which", fake_which) monkeypatch.setattr("shutil.which", fake_which)
monkeypatch.setattr("subprocess.run", _fake_run_ok) # yt-dlp probe really executes now
ch = YouTubeChannel() ch = YouTubeChannel()
status, _msg = ch.check() status, _msg = ch.check()
assert status == "ok" assert status == "ok"
assert ch.active_backend == "yt-dlp"
def test_douyin_check_does_not_call_with_invalid_url(monkeypatch, tmp_path):
"""Douyin check should use 'mcporter list' instead of calling with a hardcoded URL."""
import subprocess
from agent_reach.channels.douyin import DouyinChannel
calls = []
original_run = subprocess.run
def tracking_run(cmd, **kwargs):
calls.append(cmd)
# Simulate mcporter config list returning douyin
if "config" in cmd and "list" in cmd:
class R:
stdout = "douyin http://localhost:18070/mcp"
returncode = 0
return R()
# Simulate mcporter list douyin returning tools
if "list" in cmd and "douyin" in cmd:
class R:
stdout = "parse_douyin_video_info"
returncode = 0
return R()
return original_run(cmd, **kwargs)
monkeypatch.setattr(
"shutil.which", lambda cmd: "/usr/bin/mcporter" if cmd == "mcporter" else None
)
monkeypatch.setattr("subprocess.run", tracking_run)
ch = DouyinChannel()
status, _msg = ch.check()
# Should NOT contain any hardcoded douyin.com URL in subprocess calls
for call in calls:
call_str = " ".join(call) if isinstance(call, list) else str(call)
assert "https://www.douyin.com" not in call_str
def test_channel_can_handle_contract(): def test_channel_can_handle_contract():
@@ -171,7 +141,9 @@ def test_channel_can_handle_contract():
"reddit": "https://reddit.com/r/python", "reddit": "https://reddit.com/r/python",
"bilibili": "https://www.bilibili.com/video/BV1xx411", "bilibili": "https://www.bilibili.com/video/BV1xx411",
"xiaohongshu": "https://www.xiaohongshu.com/explore/123", "xiaohongshu": "https://www.xiaohongshu.com/explore/123",
"douyin": "https://www.douyin.com/video/123",
"linkedin": "https://www.linkedin.com/in/test", "linkedin": "https://www.linkedin.com/in/test",
"weibo": "https://weibo.com/u/1749127163",
"rss": "https://example.com/feed.xml", "rss": "https://example.com/feed.xml",
"xueqiu": "https://xueqiu.com/S/SH600519", "xueqiu": "https://xueqiu.com/S/SH600519",
"exa_search": "https://example.com", "exa_search": "https://example.com",
+11 -416
View File
@@ -654,8 +654,6 @@ class TestRedditChannel:
assert status == "off" assert status == "off"
assert "rdt-cli" in msg assert "rdt-cli" in msg
assert "public-clis/rdt-cli" in msg assert "public-clis/rdt-cli" in msg
assert "git+https://github.com/public-clis/rdt-cli.git" in msg
assert "rdt-cli>=0.4.2" not in msg
def test_reports_ok_when_authenticated(self, monkeypatch): def test_reports_ok_when_authenticated(self, monkeypatch):
monkeypatch.setattr(shutil, "which", lambda _: "/usr/local/bin/rdt") monkeypatch.setattr(shutil, "which", lambda _: "/usr/local/bin/rdt")
@@ -670,11 +668,9 @@ class TestRedditChannel:
monkeypatch.setattr(subprocess, "run", fake_run) monkeypatch.setattr(subprocess, "run", fake_run)
from agent_reach.channels.reddit import RedditChannel from agent_reach.channels.reddit import RedditChannel
ch = RedditChannel() status, msg = RedditChannel().check()
status, msg = ch.check()
assert status == "ok" assert status == "ok"
assert "testuser" in msg assert "testuser" in msg
assert ch.active_backend == "rdt-cli"
def test_reports_warn_when_not_authenticated(self, monkeypatch): def test_reports_warn_when_not_authenticated(self, monkeypatch):
monkeypatch.setattr(shutil, "which", lambda _: "/usr/local/bin/rdt") monkeypatch.setattr(shutil, "which", lambda _: "/usr/local/bin/rdt")
@@ -689,18 +685,14 @@ class TestRedditChannel:
monkeypatch.setattr(subprocess, "run", fake_run) monkeypatch.setattr(subprocess, "run", fake_run)
from agent_reach.channels.reddit import RedditChannel from agent_reach.channels.reddit import RedditChannel
ch = RedditChannel() status, msg = RedditChannel().check()
status, msg = ch.check()
assert status == "warn" assert status == "warn"
assert "403" in msg assert "403" in msg
assert "rdt login" in msg assert "rdt login" in msg
assert "Cookie-Editor" in msg assert "Cookie-Editor" in msg
assert "chromewebstore.google.com" in msg assert "chromewebstore.google.com" in msg
# 未登录是业务态:进程活着,后端仍然算可用
assert ch.active_backend == "rdt-cli"
def test_reports_error_when_status_check_fails(self, monkeypatch): def test_reports_warn_when_status_check_fails(self, monkeypatch):
"""rdt 非零退出且输出不可解析 → 工具异常(error),不再算 warn。"""
monkeypatch.setattr(shutil, "which", lambda _: "/usr/local/bin/rdt") monkeypatch.setattr(shutil, "which", lambda _: "/usr/local/bin/rdt")
def fake_run(cmd, **kwargs): def fake_run(cmd, **kwargs):
@@ -708,43 +700,8 @@ class TestRedditChannel:
monkeypatch.setattr(subprocess, "run", fake_run) monkeypatch.setattr(subprocess, "run", fake_run)
from agent_reach.channels.reddit import RedditChannel from agent_reach.channels.reddit import RedditChannel
ch = RedditChannel() status, msg = RedditChannel().check()
status, msg = ch.check() assert status == "warn"
assert status == "error"
assert "rdt 异常退出" in msg
assert ch.active_backend is None
def test_reports_error_with_reinstall_hint_when_broken(self, monkeypatch):
"""which 命中但 exec 抛 FileNotFoundErrorvenv 断链)→ error + 重装处方。"""
monkeypatch.setattr(shutil, "which", lambda _: "/usr/local/bin/rdt")
def fake_run(cmd, **kwargs):
raise FileNotFoundError("/usr/local/bin/rdt")
monkeypatch.setattr(subprocess, "run", fake_run)
from agent_reach.channels.reddit import RedditChannel
ch = RedditChannel()
status, msg = ch.check()
assert status == "error"
assert "无法执行" in msg
assert "pipx install --force" in msg # rdt 专用 git 源重装处方
assert "git+https://github.com/public-clis/rdt-cli.git" in msg
assert ch.active_backend is None
def test_reports_error_with_reinstall_hint_on_exit_127(self, monkeypatch):
"""退出码 127(找到但跑不动)同样按断链处理。"""
monkeypatch.setattr(shutil, "which", lambda _: "/usr/local/bin/rdt")
def fake_run(cmd, **kwargs):
return subprocess.CompletedProcess(cmd, 127, "", "")
monkeypatch.setattr(subprocess, "run", fake_run)
from agent_reach.channels.reddit import RedditChannel
ch = RedditChannel()
status, msg = ch.check()
assert status == "error"
assert "pipx install --force" in msg
assert ch.active_backend is None
def test_can_handle_reddit_urls(self): def test_can_handle_reddit_urls(self):
from agent_reach.channels.reddit import RedditChannel from agent_reach.channels.reddit import RedditChannel
@@ -756,25 +713,7 @@ class TestRedditChannel:
class TestXiaoHongShuChannel: class TestXiaoHongShuChannel:
"""多后端选择逻辑:OpenCLI > xiaohongshu-mcp > xhs-cli,第一个完整可用者获胜。"""
@staticmethod
def _isolate(monkeypatch, opencli=None, mcp_reachable=False):
"""隔离 OpenCLI / mcp 候选,让测试聚焦目标后端。
opencli: None 表示未安装否则传入 (status, message) 二元组
"""
import agent_reach.channels.xiaohongshu as xhs_mod
monkeypatch.setattr(
XiaoHongShuChannel, "_check_opencli", lambda self: opencli
)
monkeypatch.setattr(
xhs_mod, "_mcp_service_reachable", lambda timeout=3: mcp_reachable
)
def test_reports_ok_when_cli_authenticated(self, monkeypatch): def test_reports_ok_when_cli_authenticated(self, monkeypatch):
self._isolate(monkeypatch)
monkeypatch.setattr(shutil, "which", lambda _: "/usr/local/bin/xhs") monkeypatch.setattr(shutil, "which", lambda _: "/usr/local/bin/xhs")
def fake_run(cmd, **kwargs): def fake_run(cmd, **kwargs):
@@ -782,14 +721,11 @@ class TestXiaoHongShuChannel:
monkeypatch.setattr(subprocess, "run", fake_run) monkeypatch.setattr(subprocess, "run", fake_run)
ch = XiaoHongShuChannel() status, msg = XiaoHongShuChannel().check()
status, msg = ch.check()
assert status == "ok" assert status == "ok"
assert "xhs-cli 可用" in msg assert "完整可用" in msg
assert ch.active_backend == "xhs-cli (xiaohongshu-cli)"
def test_reports_warn_when_not_authenticated(self, monkeypatch): def test_reports_warn_when_not_authenticated(self, monkeypatch):
self._isolate(monkeypatch)
monkeypatch.setattr(shutil, "which", lambda _: "/usr/local/bin/xhs") monkeypatch.setattr(shutil, "which", lambda _: "/usr/local/bin/xhs")
def fake_run(cmd, **kwargs): def fake_run(cmd, **kwargs):
@@ -797,353 +733,12 @@ class TestXiaoHongShuChannel:
monkeypatch.setattr(subprocess, "run", fake_run) monkeypatch.setattr(subprocess, "run", fake_run)
ch = XiaoHongShuChannel() status, msg = XiaoHongShuChannel().check()
status, msg = ch.check()
assert status == "warn" assert status == "warn"
assert "xhs login" in msg assert "xhs login" in msg
# 未登录是业务态:工具进程活着,后端仍可用
assert ch.active_backend == "xhs-cli (xiaohongshu-cli)"
def test_reports_off_when_nothing_installed(self, monkeypatch): def test_reports_off_when_not_installed(self, monkeypatch):
self._isolate(monkeypatch)
monkeypatch.setattr(shutil, "which", lambda _: None) monkeypatch.setattr(shutil, "which", lambda _: None)
ch = XiaoHongShuChannel() status, msg = XiaoHongShuChannel().check()
status, msg = ch.check()
assert status == "off" assert status == "off"
# off 指引推荐当代后端,而非停更的 xhs-cli assert "xiaohongshu-cli" in msg
assert "opencli" in msg
assert "xiaohongshu-mcp" in msg
assert ch.active_backend is None
def test_reports_error_with_reinstall_hint_when_broken(self, monkeypatch):
"""which 命中但 exec 抛 FileNotFoundErrorvenv 断链)→ error + 重装处方。"""
self._isolate(monkeypatch)
monkeypatch.setattr(shutil, "which", lambda _: "/usr/local/bin/xhs")
def fake_run(cmd, **kwargs):
raise FileNotFoundError("/usr/local/bin/xhs")
monkeypatch.setattr(subprocess, "run", fake_run)
ch = XiaoHongShuChannel()
status, msg = ch.check()
assert status == "error"
assert "无法执行" in msg
assert "uv tool install --force xiaohongshu-cli" in msg
assert "pipx reinstall xiaohongshu-cli" in msg
assert ch.active_backend is None
def test_opencli_ready_wins_over_cli(self, monkeypatch):
"""OpenCLI 完整可用时按序获胜,即使 xhs-cli 也完整可用。"""
self._isolate(monkeypatch, opencli=("ok", "OpenCLI 可用(复用浏览器登录态)"))
monkeypatch.setattr(shutil, "which", lambda _: "/usr/local/bin/xhs")
def fake_run(cmd, **kwargs):
return subprocess.CompletedProcess(cmd, 0, "ok: true\n", "")
monkeypatch.setattr(subprocess, "run", fake_run)
ch = XiaoHongShuChannel()
status, msg = ch.check()
assert status == "ok"
assert ch.active_backend == "OpenCLI"
def test_opencli_warn_loses_to_usable_cli(self, monkeypatch):
"""OpenCLI 装了但扩展未连(warn)时,完整可用的 xhs-cli 获胜。"""
self._isolate(monkeypatch, opencli=("warn", "扩展未连接"))
monkeypatch.setattr(shutil, "which", lambda _: "/usr/local/bin/xhs")
def fake_run(cmd, **kwargs):
return subprocess.CompletedProcess(cmd, 0, "ok: true\n", "")
monkeypatch.setattr(subprocess, "run", fake_run)
ch = XiaoHongShuChannel()
status, msg = ch.check()
assert status == "ok"
assert ch.active_backend == "xhs-cli (xiaohongshu-cli)"
def test_mcp_service_wins_when_opencli_absent(self, monkeypatch):
"""服务器场景:OpenCLI 未装、mcp 服务可达且 mcporter 已接入 → mcp 获胜。"""
self._isolate(monkeypatch, mcp_reachable=True)
def fake_which(name):
return "/usr/local/bin/mcporter" if name == "mcporter" else None
monkeypatch.setattr(shutil, "which", fake_which)
def fake_run(cmd, **kwargs):
return subprocess.CompletedProcess(cmd, 0, "exa\nxiaohongshu\n", "")
monkeypatch.setattr(subprocess, "run", fake_run)
ch = XiaoHongShuChannel()
status, msg = ch.check()
assert status == "ok"
assert ch.active_backend == "xiaohongshu-mcp"
assert "search_feeds" in msg
def test_mcp_reachable_but_mcporter_unconfigured_warns(self, monkeypatch):
self._isolate(monkeypatch, mcp_reachable=True)
def fake_which(name):
return "/usr/local/bin/mcporter" if name == "mcporter" else None
monkeypatch.setattr(shutil, "which", fake_which)
def fake_run(cmd, **kwargs):
return subprocess.CompletedProcess(cmd, 0, "exa\n", "")
monkeypatch.setattr(subprocess, "run", fake_run)
ch = XiaoHongShuChannel()
status, msg = ch.check()
assert status == "warn"
assert "mcporter config add xiaohongshu" in msg
assert ch.active_backend == "xiaohongshu-mcp"
def test_backend_override_prefers_cli(self, monkeypatch):
"""config xiaohongshu_backend=xhs-cli 时,即使 OpenCLI ready 也用 xhs-cli。"""
self._isolate(monkeypatch, opencli=("ok", "OpenCLI 可用"))
monkeypatch.setattr(shutil, "which", lambda _: "/usr/local/bin/xhs")
def fake_run(cmd, **kwargs):
return subprocess.CompletedProcess(cmd, 0, "ok: true\n", "")
monkeypatch.setattr(subprocess, "run", fake_run)
class _Cfg:
def get(self, key, default=None):
return "xhs-cli" if key == "xiaohongshu_backend" else default
ch = XiaoHongShuChannel()
status, _ = ch.check(_Cfg())
assert status == "ok"
assert ch.active_backend == "xhs-cli (xiaohongshu-cli)"
class TestBilibiliChannel:
def test_reports_error_with_reinstall_hint_when_ytdlp_broken(self, monkeypatch):
"""yt-dlp which 命中但 exec 失败(venv 断链)→ error + 重装处方。"""
monkeypatch.setattr(shutil, "which", lambda _: "/usr/local/bin/yt-dlp")
def fake_run(cmd, **kwargs):
raise FileNotFoundError(cmd[0])
monkeypatch.setattr(subprocess, "run", fake_run)
from agent_reach.channels.bilibili import BilibiliChannel
ch = BilibiliChannel()
status, msg = ch.check()
assert status == "error"
assert "无法执行" in msg
assert "uv tool install --force yt-dlp" in msg
assert "pipx reinstall yt-dlp" in msg
assert ch.active_backend is None
def test_active_backend_set_when_ytdlp_and_bili_ok(self, monkeypatch):
monkeypatch.setattr(
shutil, "which",
lambda cmd: f"/usr/local/bin/{cmd}" if cmd in ("yt-dlp", "bili") else None,
)
def fake_run(cmd, **kwargs):
return subprocess.CompletedProcess(cmd, 0, "2026.06.09", "")
monkeypatch.setattr(subprocess, "run", fake_run)
from agent_reach.channels.bilibili import BilibiliChannel
ch = BilibiliChannel()
status, msg = ch.check()
assert status == "ok"
assert "bili-cli 可用" in msg
assert ch.active_backend == "yt-dlp"
def test_bili_broken_does_not_count_as_available(self, monkeypatch):
"""bili-cli 断链时不计为可用,降级走搜索 APIyt-dlp 仍是 active_backend。"""
monkeypatch.setattr(
shutil, "which",
lambda cmd: f"/usr/local/bin/{cmd}" if cmd in ("yt-dlp", "bili") else None,
)
def fake_run(cmd, **kwargs):
if "yt-dlp" in cmd[0]:
return subprocess.CompletedProcess(cmd, 0, "2026.06.09", "")
raise FileNotFoundError(cmd[0])
monkeypatch.setattr(subprocess, "run", fake_run)
import agent_reach.channels.bilibili as bilibili_mod
monkeypatch.setattr(bilibili_mod, "_search_api_ok", lambda: True)
ch = bilibili_mod.BilibiliChannel()
status, msg = ch.check()
assert status == "ok" # 搜索 API 兜底
assert "不计为可用" in msg
assert ch.active_backend == "yt-dlp"
class TestYouTubeChannel:
def test_reports_error_with_reinstall_hint_when_broken(self, monkeypatch):
"""yt-dlp which 命中但 exec 抛 FileNotFoundError → error + 重装处方。"""
monkeypatch.setattr(shutil, "which", lambda _: "/usr/local/bin/yt-dlp")
def fake_run(cmd, **kwargs):
raise FileNotFoundError(cmd[0])
monkeypatch.setattr(subprocess, "run", fake_run)
from agent_reach.channels.youtube import YouTubeChannel
ch = YouTubeChannel()
status, msg = ch.check()
assert status == "error"
assert "无法执行" in msg
assert "uv tool install --force yt-dlp" in msg
assert ch.active_backend is None
class TestGitHubChannel:
def test_reports_error_with_reinstall_hint_when_broken(self, monkeypatch):
"""gh which 命中但 exec 失败 → error + brew 重装处方(gh 不是 pip 包)。"""
monkeypatch.setattr(shutil, "which", lambda _: "/usr/local/bin/gh")
def fake_run(cmd, **kwargs):
raise FileNotFoundError(cmd[0])
monkeypatch.setattr(subprocess, "run", fake_run)
from agent_reach.channels.github import GitHubChannel
ch = GitHubChannel()
status, msg = ch.check()
assert status == "error"
assert "无法执行" in msg
assert "brew reinstall gh" in msg
assert ch.active_backend is None
def test_active_backend_set_when_authenticated(self, monkeypatch):
monkeypatch.setattr(shutil, "which", lambda _: "/usr/local/bin/gh")
def fake_run(cmd, **kwargs):
return subprocess.CompletedProcess(cmd, 0, "Logged in to github.com", "")
monkeypatch.setattr(subprocess, "run", fake_run)
from agent_reach.channels.github import GitHubChannel
ch = GitHubChannel()
status, msg = ch.check()
assert status == "ok"
assert ch.active_backend == "gh CLI"
def test_active_backend_set_when_unauthenticated(self, monkeypatch):
"""gh auth status 非零退出是正常业务态(未登录):warn 但后端可用。"""
monkeypatch.setattr(shutil, "which", lambda _: "/usr/local/bin/gh")
def fake_run(cmd, **kwargs):
return subprocess.CompletedProcess(cmd, 1, "", "You are not logged in")
monkeypatch.setattr(subprocess, "run", fake_run)
from agent_reach.channels.github import GitHubChannel
ch = GitHubChannel()
status, msg = ch.check()
assert status == "warn"
assert "gh auth login" in msg
assert ch.active_backend == "gh CLI"
class TestLinkedInChannel:
def test_reports_error_with_reinstall_hint_when_broken(self, monkeypatch):
"""mcporter which 命中但 exec 失败 → error + npm 重装处方。"""
monkeypatch.setattr(shutil, "which", lambda _: "/usr/local/bin/mcporter")
def fake_run(cmd, **kwargs):
raise FileNotFoundError(cmd[0])
monkeypatch.setattr(subprocess, "run", fake_run)
from agent_reach.channels.linkedin import LinkedInChannel
ch = LinkedInChannel()
status, msg = ch.check()
assert status == "error"
assert "npm install -g mcporter" in msg
assert ch.active_backend is None
def test_active_backend_set_when_linkedin_configured(self, monkeypatch):
monkeypatch.setattr(shutil, "which", lambda _: "/usr/local/bin/mcporter")
def fake_run(cmd, **kwargs):
return subprocess.CompletedProcess(cmd, 0, "linkedin http://localhost:3000/mcp", "")
monkeypatch.setattr(subprocess, "run", fake_run)
from agent_reach.channels.linkedin import LinkedInChannel
ch = LinkedInChannel()
status, msg = ch.check()
assert status == "ok"
assert ch.active_backend == "linkedin-scraper-mcp"
def test_off_without_backend_when_linkedin_not_configured(self, monkeypatch):
monkeypatch.setattr(shutil, "which", lambda _: "/usr/local/bin/mcporter")
def fake_run(cmd, **kwargs):
return subprocess.CompletedProcess(cmd, 0, "exa https://mcp.exa.ai/mcp", "")
monkeypatch.setattr(subprocess, "run", fake_run)
from agent_reach.channels.linkedin import LinkedInChannel
ch = LinkedInChannel()
status, msg = ch.check()
assert status == "off"
assert ch.active_backend is None
class TestExaSearchChannel:
def test_reports_error_with_reinstall_hint_when_broken(self, monkeypatch):
"""mcporter which 命中但 exec 失败 → error + npm 重装处方。"""
monkeypatch.setattr(shutil, "which", lambda _: "/usr/local/bin/mcporter")
def fake_run(cmd, **kwargs):
raise FileNotFoundError(cmd[0])
monkeypatch.setattr(subprocess, "run", fake_run)
from agent_reach.channels.exa_search import ExaSearchChannel
ch = ExaSearchChannel()
status, msg = ch.check()
assert status == "error"
assert "npm install -g mcporter" in msg
assert ch.active_backend is None
def test_active_backend_set_when_exa_configured(self, monkeypatch):
monkeypatch.setattr(shutil, "which", lambda _: "/usr/local/bin/mcporter")
def fake_run(cmd, **kwargs):
return subprocess.CompletedProcess(cmd, 0, "exa https://mcp.exa.ai/mcp", "")
monkeypatch.setattr(subprocess, "run", fake_run)
from agent_reach.channels.exa_search import ExaSearchChannel
ch = ExaSearchChannel()
status, msg = ch.check()
assert status == "ok"
assert ch.active_backend == "Exa via mcporter"
class TestXiaoyuzhouChannel:
def test_reports_error_with_reinstall_hint_when_ffmpeg_broken(self, monkeypatch):
"""ffmpeg which 命中但 exec 失败(pip 假 ffmpeg 断链)→ error + 重装处方。"""
monkeypatch.setattr(shutil, "which", lambda _: "/usr/local/bin/ffmpeg")
def fake_run(cmd, **kwargs):
raise FileNotFoundError(cmd[0])
monkeypatch.setattr(subprocess, "run", fake_run)
from agent_reach.channels.xiaoyuzhou import XiaoyuzhouChannel
ch = XiaoyuzhouChannel()
status, msg = ch.check()
assert status == "error"
assert "无法执行" in msg
assert "brew install ffmpeg" in msg
assert ch.active_backend is None
def test_active_backend_set_when_fully_configured(self, monkeypatch):
monkeypatch.setattr(shutil, "which", lambda _: "/usr/local/bin/ffmpeg")
def fake_run(cmd, **kwargs):
return subprocess.CompletedProcess(cmd, 0, "ffmpeg version 7.0", "")
monkeypatch.setattr(subprocess, "run", fake_run)
monkeypatch.setattr("os.path.isfile", lambda p: True) # transcribe.sh 已安装
monkeypatch.setenv("GROQ_API_KEY", "gsk_test")
from agent_reach.channels.xiaoyuzhou import XiaoyuzhouChannel
ch = XiaoyuzhouChannel()
status, msg = ch.check()
assert status == "ok"
assert ch.active_backend == "groq-whisper"
+18 -40
View File
@@ -1,12 +1,12 @@
# -*- coding: utf-8 -*- # -*- coding: utf-8 -*-
"""Tests for Agent Reach CLI.""" """Tests for Agent Reach CLI."""
import shutil import json
import subprocess import os
from unittest.mock import patch
import pytest import pytest
import requests import requests
from unittest.mock import patch
import agent_reach.cli as cli import agent_reach.cli as cli
from agent_reach.cli import main from agent_reach.cli import main
@@ -33,21 +33,6 @@ class TestCLI:
assert "Agent Reach" in captured.out assert "Agent Reach" in captured.out
assert "" in captured.out assert "" in captured.out
def test_transcribe_command_prints_text(self, capsys):
with patch("agent_reach.transcribe.transcribe", return_value="hello transcript"):
with patch("sys.argv", ["agent-reach", "transcribe", "audio.mp3"]):
main()
captured = capsys.readouterr()
assert "hello transcript" in captured.out
def test_transcribe_command_writes_output_file(self, capsys, tmp_path):
out_file = tmp_path / "t.txt"
with patch("agent_reach.transcribe.transcribe", return_value="saved text"):
with patch("sys.argv", ["agent-reach", "transcribe", "audio.mp3", "-o", str(out_file)]):
main()
assert out_file.read_text(encoding="utf-8").strip() == "saved text"
assert "Transcript written" in capsys.readouterr().out
def test_parse_twitter_cookie_input_separate_values(self): def test_parse_twitter_cookie_input_separate_values(self):
auth_token, ct0 = cli._parse_twitter_cookie_input("token123 ct0abc") auth_token, ct0 = cli._parse_twitter_cookie_input("token123 ct0abc")
assert auth_token == "token123" assert auth_token == "token123"
@@ -60,30 +45,23 @@ class TestCLI:
assert auth_token == "token123" assert auth_token == "token123"
assert ct0 == "ct0abc" assert ct0 == "ct0abc"
def test_install_reddit_deps_prefers_github_source(self, monkeypatch, capsys): def test_configure_xhs_cookies_writes_xhs_cli_cookie_file(self, tmp_path, monkeypatch, capsys):
state = {"rdt_installed": False} monkeypatch.setenv("HOME", str(tmp_path))
commands = [] monkeypatch.setattr("shutil.which", lambda _name: None)
def fake_which(name): cli._configure_xhs_cookies("a1=token123; web_session=session456; other=value")
if name == "rdt":
return "/usr/local/bin/rdt" if state["rdt_installed"] else None
if name == "pipx":
return "/usr/local/bin/pipx"
return None
def fake_run(cmd, **kwargs): cookie_path = tmp_path / ".xiaohongshu-cli" / "cookies.json"
commands.append(cmd) data = json.loads(cookie_path.read_text())
state["rdt_installed"] = True assert data["a1"] == "token123"
return subprocess.CompletedProcess(cmd, 0, "", "") assert data["web_session"] == "session456"
assert "saved_at" in data
assert oct(os.stat(cookie_path).st_mode & 0o777) == "0o600"
monkeypatch.setattr(shutil, "which", fake_which) legacy_path = tmp_path / ".agent-reach" / "xhs-cookies.json"
monkeypatch.setattr(subprocess, "run", fake_run) assert legacy_path.exists()
captured = capsys.readouterr()
cli._install_reddit_deps() assert "xhs-cli cookies saved" in captured.out
out = capsys.readouterr().out
assert commands == [["pipx", "install", cli._RDT_GIT_SOURCE]]
assert "✅ rdt-cli installed" in out
class TestCheckUpdateRetry: class TestCheckUpdateRetry:
@@ -132,7 +110,7 @@ class TestCheckUpdateRetry:
sequence = [ sequence = [
R(429, headers={"Retry-After": "3"}), R(429, headers={"Retry-After": "3"}),
R(200, payload={"tag_name": "v1.4.2"}), R(200, payload={"tag_name": "v1.4.0"}),
] ]
with patch("requests.get", side_effect=sequence): with patch("requests.get", side_effect=sequence):
-89
View File
@@ -1,89 +0,0 @@
# -*- coding: utf-8 -*-
"""Verify credential files written by cookie sync helpers and CLI helpers
are owner-only (0o600) and that values containing shell metacharacters do
not break the shell-sourceable env file produced by _sync_bird_env().
Companion to tests/test_config.py::test_save_creates_file_with_restricted_permissions
the same threat-model claim ("Cookie/Token only stored locally, 600
permissions") covers these paths.
"""
import json
import os
import stat
import subprocess
import sys
import tempfile
import pytest
from agent_reach.cookie_extract import _sync_bird_env, _sync_xfetch_session
def _owner_only(path: str) -> bool:
mode = os.stat(path).st_mode
return not (mode & (stat.S_IRGRP | stat.S_IWGRP | stat.S_IROTH | stat.S_IWOTH))
@pytest.mark.skipif(sys.platform == "win32", reason="POSIX perm semantics only")
def test_sync_xfetch_session_writes_0600(tmp_path, monkeypatch):
monkeypatch.setenv("HOME", str(tmp_path))
_sync_xfetch_session("auth_xxx", "ct0_yyy")
session_path = tmp_path / ".config" / "xfetch" / "session.json"
assert session_path.exists(), "expected ~/.config/xfetch/session.json"
assert _owner_only(str(session_path)), "session.json must be 0o600"
# Round-trip the content so we know we didn't accidentally corrupt JSON.
data = json.loads(session_path.read_text(encoding="utf-8"))
assert data["authToken"] == "auth_xxx"
assert data["ct0"] == "ct0_yyy"
@pytest.mark.skipif(sys.platform == "win32", reason="POSIX perm semantics only")
def test_sync_bird_env_writes_0600(tmp_path, monkeypatch):
monkeypatch.setenv("HOME", str(tmp_path))
_sync_bird_env("auth_xxx", "ct0_yyy")
env_path = tmp_path / ".config" / "bird" / "credentials.env"
assert env_path.exists(), "expected ~/.config/bird/credentials.env"
assert _owner_only(str(env_path)), "credentials.env must be 0o600"
@pytest.mark.skipif(sys.platform == "win32", reason="POSIX sh needed for sourcing")
def test_sync_bird_env_quotes_shell_metachars(tmp_path, monkeypatch):
"""Tokens containing ", $, `, ; etc. must not break out of the assignment.
Prior implementation used `f'AUTH_TOKEN="{auth_token}"'` which an attacker-
controlled cookie containing a literal `"` could break out of, turning a
later `source ~/.config/bird/credentials.env` into arbitrary shell.
"""
monkeypatch.setenv("HOME", str(tmp_path))
# Side-effect markers live under tmp_path (auto-cleaned by pytest) rather
# than a shared absolute /tmp path — otherwise one vulnerable run leaves a
# marker behind that fails every later run on the same machine/CI runner.
pwn_auth = tmp_path / "pwn-auth"
pwn_ct0 = tmp_path / "pwn-ct0"
hostile_auth = f'inj"; touch {pwn_auth}; #'
hostile_ct0 = f"ct0_$(touch {pwn_ct0})"
_sync_bird_env(hostile_auth, hostile_ct0)
env_path = tmp_path / ".config" / "bird" / "credentials.env"
# Sourcing the file must NOT execute the injected payload. Read back the
# exported values from a subshell instead — they should equal the originals.
probe = (
f". {env_path}; "
f'printf "AUTH=%s\\nCT0=%s\\n" "$AUTH_TOKEN" "$CT0"'
)
result = subprocess.run(
["sh", "-c", probe],
capture_output=True,
text=True,
timeout=5,
)
assert result.returncode == 0, result.stderr
lines = dict(
line.split("=", 1) for line in result.stdout.strip().splitlines() if "=" in line
)
assert lines["AUTH"] == hostile_auth, "auth_token round-trip broke — injection possible"
assert lines["CT0"] == hostile_ct0, "ct0 round-trip broke — injection possible"
# And no side-effect files materialised.
assert not pwn_auth.exists()
assert not pwn_ct0.exists()
+2 -8
View File
@@ -8,15 +8,13 @@ from agent_reach.config import Config
class _StubChannel: class _StubChannel:
def __init__(self, name, description, tier, status, message, backends=None, def __init__(self, name, description, tier, status, message, backends=None):
active_backend=None):
self.name = name self.name = name
self.description = description self.description = description
self.tier = tier self.tier = tier
self._status = status self._status = status
self._message = message self._message = message
self.backends = backends or [] self.backends = backends or []
self.active_backend = active_backend
def check(self, config=None): def check(self, config=None):
return self._status, self._message return self._status, self._message
@@ -33,8 +31,7 @@ class TestDoctor:
doctor, doctor,
"get_all_channels", "get_all_channels",
lambda: [ lambda: [
_StubChannel("web", "网页", 0, "ok", "可抓取网页", ["requests"], _StubChannel("web", "网页", 0, "ok", "可抓取网页", ["requests"]),
active_backend="requests"),
_StubChannel("github", "GitHub", 0, "warn", "gh 未安装", ["gh"]), _StubChannel("github", "GitHub", 0, "warn", "gh 未安装", ["gh"]),
_StubChannel("exa_search", "全网语义搜索", 1, "off", "mcporter 未配置", ["Exa"]), _StubChannel("exa_search", "全网语义搜索", 1, "off", "mcporter 未配置", ["Exa"]),
], ],
@@ -49,7 +46,6 @@ class TestDoctor:
"message": "可抓取网页", "message": "可抓取网页",
"tier": 0, "tier": 0,
"backends": ["requests"], "backends": ["requests"],
"active_backend": "requests",
}, },
"github": { "github": {
"status": "warn", "status": "warn",
@@ -57,7 +53,6 @@ class TestDoctor:
"message": "gh 未安装", "message": "gh 未安装",
"tier": 0, "tier": 0,
"backends": ["gh"], "backends": ["gh"],
"active_backend": None,
}, },
"exa_search": { "exa_search": {
"status": "off", "status": "off",
@@ -65,7 +60,6 @@ class TestDoctor:
"message": "mcporter 未配置", "message": "mcporter 未配置",
"tier": 1, "tier": 1,
"backends": ["Exa"], "backends": ["Exa"],
"active_backend": None,
}, },
} }
-100
View File
@@ -1,100 +0,0 @@
# -*- coding: utf-8 -*-
"""Tests for the OpenCLI cross-channel backend probing."""
from unittest.mock import patch
from agent_reach.backends import opencli_status, opencli_summary
from agent_reach.probe import ProbeResult
def _status_with(version_probe, daemon_probe=None, ext_on_disk=False):
"""Run opencli_status with probe_command and disk check patched."""
calls = []
def fake_probe(cmd, args=("--version",), **kwargs):
calls.append(list(args))
if list(args) == ["--version"]:
return version_probe
return daemon_probe or ProbeResult("missing")
with patch("agent_reach.backends.opencli.probe_command", side_effect=fake_probe), \
patch(
"agent_reach.backends.opencli._extension_installed_on_disk",
return_value=ext_on_disk,
):
return opencli_status(), calls
def test_not_installed():
st, _ = _status_with(ProbeResult("missing"))
assert not st.installed
assert not st.ready
assert "未安装" in opencli_summary(st)
def test_broken_node_env_gives_npm_hint():
st, _ = _status_with(ProbeResult("broken", hint="x"))
assert st.installed and st.broken
assert "npm install -g @jackwener/opencli" in st.hint
assert not st.ready
def test_daemon_running_extension_connected_is_ready():
daemon_out = "Daemon: running (PID 37389)\nVersion: v1.8.3\nExtension: connected\n"
st, _ = _status_with(
ProbeResult("ok", output="1.8.3"),
ProbeResult("ok", output=daemon_out),
)
assert st.installed and st.daemon_running and st.extension_connected
assert st.ready
assert "1.8.3" in opencli_summary(st)
def test_extension_never_installed_not_ready_with_store_guide():
daemon_out = "Daemon: running (PID 1)\nExtension: disconnected\n"
st, _ = _status_with(
ProbeResult("ok", output="1.8.3"),
ProbeResult("ok", output=daemon_out),
ext_on_disk=False,
)
assert st.daemon_running and not st.extension_connected
assert not st.ready
assert "chromewebstore.google.com" in st.hint
def test_sleeping_extension_counts_as_ready():
"""实测:扩展 service worker 睡眠时 daemon status 报 disconnected,
但任何真实命令会唤醒它装在磁盘上即视为可用"""
daemon_out = "Daemon: running (PID 1)\nExtension: disconnected\n"
st, _ = _status_with(
ProbeResult("ok", output="1.8.3"),
ProbeResult("ok", output=daemon_out),
ext_on_disk=True,
)
assert not st.extension_connected
assert st.extension_installed
assert st.ready
assert "唤醒" in opencli_summary(st)
assert st.hint == ""
def test_daemon_not_running_parsed_correctly():
st, _ = _status_with(
ProbeResult("ok", output="1.8.3"),
ProbeResult("ok", output="Daemon: not running\n"),
)
assert st.installed
assert not st.daemon_running
assert not st.extension_connected
assert "自动启动" in opencli_summary(st)
def test_probe_uses_daemon_status_not_doctor():
"""`opencli doctor` auto-starts the daemon (side effect) — health checks
must only ever call `daemon status`."""
_, calls = _status_with(
ProbeResult("ok", output="1.8.3"),
ProbeResult("ok", output="Daemon: not running\n"),
)
assert ["daemon", "status"] in calls
assert ["doctor"] not in calls
-90
View File
@@ -1,90 +0,0 @@
# -*- coding: utf-8 -*-
"""Tests for agent_reach.probe — real-execution probing and failure classification."""
import os
import stat
import sys
import pytest
from agent_reach.probe import ProbeResult, probe_command, reinstall_hint
def _make_executable(path, content):
path.write_text(content)
path.chmod(path.stat().st_mode | stat.S_IXUSR | stat.S_IXGRP | stat.S_IXOTH)
return str(path)
def test_missing_command():
r = probe_command("definitely-not-a-real-command-xyz")
assert r.status == "missing"
assert not r.ok
@pytest.mark.skipif(sys.platform == "win32", reason="shebang semantics are POSIX-only")
def test_broken_shebang_detected_as_broken(tmp_path, monkeypatch):
"""A stale venv shim: which() finds it, exec raises FileNotFoundError."""
script = _make_executable(
tmp_path / "stale-tool", "#!/nonexistent/venv/bin/python\nprint('hi')\n"
)
monkeypatch.setenv("PATH", str(tmp_path) + os.pathsep + os.environ.get("PATH", ""))
r = probe_command("stale-tool", package="stale-tool-pkg")
assert r.status == "broken"
assert "uv tool install --force stale-tool-pkg" in r.hint
assert "pipx reinstall stale-tool-pkg" in r.hint
@pytest.mark.skipif(sys.platform == "win32", reason="shell script fixture is POSIX-only")
def test_healthy_command_returns_ok_with_output(tmp_path, monkeypatch):
script = _make_executable(
tmp_path / "healthy-tool", "#!/bin/sh\necho 'healthy-tool 1.2.3'\n"
)
monkeypatch.setenv("PATH", str(tmp_path) + os.pathsep + os.environ.get("PATH", ""))
r = probe_command("healthy-tool")
assert r.ok
assert "1.2.3" in r.output
@pytest.mark.skipif(sys.platform == "win32", reason="shell script fixture is POSIX-only")
def test_nonzero_exit_classified_as_error(tmp_path, monkeypatch):
script = _make_executable(
tmp_path / "failing-tool", "#!/bin/sh\necho 'boom' >&2\nexit 3\n"
)
monkeypatch.setenv("PATH", str(tmp_path) + os.pathsep + os.environ.get("PATH", ""))
r = probe_command("failing-tool")
assert r.status == "error"
assert "boom" in r.output
@pytest.mark.skipif(sys.platform == "win32", reason="shell script fixture is POSIX-only")
def test_exit_127_classified_as_broken(tmp_path, monkeypatch):
script = _make_executable(tmp_path / "wrapper-tool", "#!/bin/sh\nexit 127\n")
monkeypatch.setenv("PATH", str(tmp_path) + os.pathsep + os.environ.get("PATH", ""))
r = probe_command("wrapper-tool", package="wrapper-pkg")
assert r.status == "broken"
assert "wrapper-pkg" in r.hint
@pytest.mark.skipif(sys.platform == "win32", reason="shell script fixture is POSIX-only")
def test_retries_help_transient_failures(tmp_path, monkeypatch):
"""First call fails (exit 1), second succeeds — retries=1 should return ok."""
marker = tmp_path / "ran-once"
script = _make_executable(
tmp_path / "flaky-tool",
f"#!/bin/sh\nif [ -f {marker} ]; then echo ok; exit 0; fi\ntouch {marker}\nexit 1\n",
)
monkeypatch.setenv("PATH", str(tmp_path) + os.pathsep + os.environ.get("PATH", ""))
r = probe_command("flaky-tool", retries=1)
assert r.ok
def test_reinstall_hint_mentions_both_installers():
hint = reinstall_hint("some-pkg")
assert "uv tool install --force some-pkg" in hint
assert "pipx reinstall some-pkg" in hint
-18
View File
@@ -1,18 +0,0 @@
from agent_reach.utils.process import mcporter_utf8_env_args, utf8_subprocess_env
def test_utf8_subprocess_env_forces_python_utf8():
env = utf8_subprocess_env({"PYTHONUTF8": "0", "OTHER": "value"})
assert env["PYTHONUTF8"] == "1"
assert env["PYTHONIOENCODING"] == "utf-8"
assert env["OTHER"] == "value"
def test_mcporter_utf8_env_args():
assert mcporter_utf8_env_args() == [
"--env",
"PYTHONUTF8=1",
"--env",
"PYTHONIOENCODING=utf-8",
]
+8 -5
View File
@@ -39,11 +39,14 @@ class TestSkillCommand(unittest.TestCase):
with patch.dict(os.environ, env, clear=True): with patch.dict(os.environ, env, clear=True):
_install_skill() _install_skill()
target = os.path.join(skill_dir, "agent-reach", "SKILL.md")
# Check at least one known skill dir pattern # Check at least one known skill dir pattern
found = False
for dirpath, _, filenames in os.walk(tmpdir): for dirpath, _, filenames in os.walk(tmpdir):
if "SKILL.md" in filenames: if "SKILL.md" in filenames:
found = True
# Verify content is non-empty # Verify content is non-empty
with open(os.path.join(dirpath, "SKILL.md"), encoding="utf-8") as f: with open(os.path.join(dirpath, "SKILL.md")) as f:
content = f.read() content = f.read()
self.assertIn("Agent Reach", content) self.assertIn("Agent Reach", content)
# _install_skill may or may not find dirs depending on mock; just ensure no crash # _install_skill may or may not find dirs depending on mock; just ensure no crash
@@ -55,7 +58,7 @@ class TestSkillCommand(unittest.TestCase):
# Create a fake skill installation # Create a fake skill installation
skill_path = os.path.join(tmpdir, ".openclaw", "skills", "agent-reach") skill_path = os.path.join(tmpdir, ".openclaw", "skills", "agent-reach")
os.makedirs(skill_path) os.makedirs(skill_path)
with open(os.path.join(skill_path, "SKILL.md"), "w", encoding="utf-8") as f: with open(os.path.join(skill_path, "SKILL.md"), "w") as f:
f.write("test") f.write("test")
self.assertTrue(os.path.exists(skill_path)) self.assertTrue(os.path.exists(skill_path))
@@ -89,7 +92,7 @@ class TestSkillCommand(unittest.TestCase):
target = os.path.join(skill_parent, "agent-reach", "SKILL.md") target = os.path.join(skill_parent, "agent-reach", "SKILL.md")
self.assertTrue(os.path.exists(target)) self.assertTrue(os.path.exists(target))
with open(target, encoding="utf-8") as f: with open(target) as f:
content = f.read() content = f.read()
self.assertIn("Agent Reach", content) self.assertIn("Agent Reach", content)
@@ -111,10 +114,10 @@ class TestSkillCommand(unittest.TestCase):
target = os.path.join(skill_parent, "agent-reach", "SKILL.md") target = os.path.join(skill_parent, "agent-reach", "SKILL.md")
self.assertTrue(os.path.exists(target)) self.assertTrue(os.path.exists(target))
with open(target, encoding="utf-8") as f: with open(target) as f:
content = f.read() content = f.read()
self.assertTrue(content.strip()) self.assertTrue(content.strip())
self.assertIn("Xiaoyuzhou Podcast, LinkedIn", content) self.assertIn("Give your AI agent eyes to see the entire internet.", content)
self.assertNotIn("搜推特", content) self.assertNotIn("搜推特", content)
self.assertTrue( self.assertTrue(
os.path.exists(os.path.join(skill_parent, "agent-reach", "references")) os.path.exists(os.path.join(skill_parent, "agent-reach", "references"))
-262
View File
@@ -1,262 +0,0 @@
# -*- coding: utf-8 -*-
"""Tests for agent_reach.transcribe — provider routing, fallback, and errors."""
from typing import List
import pytest
from agent_reach import transcribe as tr
from agent_reach.config import Config
# --- Fixtures ----------------------------------------------------------- #
@pytest.fixture
def fake_config(tmp_path, monkeypatch):
"""A Config that writes to a temp dir and never touches the user's HOME."""
cfg_path = tmp_path / "config.yaml"
monkeypatch.setattr(Config, "CONFIG_DIR", tmp_path)
monkeypatch.setattr(Config, "CONFIG_FILE", cfg_path)
cfg = Config(config_path=cfg_path)
return cfg
@pytest.fixture
def chunk_file(tmp_path):
p = tmp_path / "chunk.m4a"
p.write_bytes(b"\x00fake-m4a-bytes")
return p
class FakeResponse:
def __init__(self, status_code: int, text: str = ""):
self.status_code = status_code
self.text = text
@property
def ok(self) -> bool:
return 200 <= self.status_code < 300
# --- transcribe_chunk: provider routing -------------------------------- #
class TestTranscribeChunk:
def test_routes_to_groq_endpoint(self, monkeypatch, fake_config, chunk_file):
fake_config.set("groq_api_key", "gsk_test")
captured = {}
def fake_post(url, headers=None, files=None, data=None, timeout=None):
captured["url"] = url
captured["headers"] = headers
captured["model"] = data["model"]
return FakeResponse(200, "hello world")
monkeypatch.setattr(tr.requests, "post", fake_post)
text = tr.transcribe_chunk(chunk_file, "groq", config=fake_config)
assert text == "hello world"
assert captured["url"] == tr.PROVIDERS["groq"]["endpoint"]
assert captured["model"] == "whisper-large-v3"
assert captured["headers"]["Authorization"] == "Bearer gsk_test"
def test_routes_to_openai_endpoint(self, monkeypatch, fake_config, chunk_file):
fake_config.set("openai_api_key", "sk-test")
captured = {}
def fake_post(url, headers=None, files=None, data=None, timeout=None):
captured["url"] = url
captured["model"] = data["model"]
return FakeResponse(200, "openai output")
monkeypatch.setattr(tr.requests, "post", fake_post)
text = tr.transcribe_chunk(chunk_file, "openai", config=fake_config)
assert text == "openai output"
assert captured["url"] == tr.PROVIDERS["openai"]["endpoint"]
assert captured["model"] == "whisper-1"
def test_raises_when_key_missing(self, fake_config, chunk_file):
with pytest.raises(tr.NoProviderConfigured):
tr.transcribe_chunk(chunk_file, "groq", config=fake_config)
def test_raises_on_http_error(self, monkeypatch, fake_config, chunk_file):
fake_config.set("groq_api_key", "gsk_test")
monkeypatch.setattr(
tr.requests,
"post",
lambda *a, **k: FakeResponse(429, "rate limited"),
)
with pytest.raises(tr.TranscribeError, match="HTTP 429"):
tr.transcribe_chunk(chunk_file, "groq", config=fake_config)
def test_unknown_provider(self, fake_config, chunk_file):
with pytest.raises(tr.TranscribeError, match="unknown provider"):
tr.transcribe_chunk(chunk_file, "azure", config=fake_config)
# --- _transcribe_with_fallback ----------------------------------------- #
class TestFallback:
def test_groq_succeeds_no_openai_call(self, monkeypatch, fake_config, chunk_file):
fake_config.set("groq_api_key", "gsk_test")
fake_config.set("openai_api_key", "sk-test")
calls: List[str] = []
def fake_post(url, headers=None, files=None, data=None, timeout=None):
calls.append(url)
return FakeResponse(200, "from-groq")
monkeypatch.setattr(tr.requests, "post", fake_post)
text = tr._transcribe_with_fallback(chunk_file, ["groq", "openai"], fake_config)
assert text == "from-groq"
assert calls == [tr.PROVIDERS["groq"]["endpoint"]]
def test_groq_429_falls_back_to_openai(self, monkeypatch, fake_config, chunk_file):
fake_config.set("groq_api_key", "gsk_test")
fake_config.set("openai_api_key", "sk-test")
calls: List[str] = []
def fake_post(url, headers=None, files=None, data=None, timeout=None):
calls.append(url)
if url == tr.PROVIDERS["groq"]["endpoint"]:
return FakeResponse(429, "rate limited")
return FakeResponse(200, "from-openai")
monkeypatch.setattr(tr.requests, "post", fake_post)
text = tr._transcribe_with_fallback(chunk_file, ["groq", "openai"], fake_config)
assert text == "from-openai"
assert calls == [
tr.PROVIDERS["groq"]["endpoint"],
tr.PROVIDERS["openai"]["endpoint"],
]
def test_skip_unconfigured_provider(self, monkeypatch, fake_config, chunk_file):
# Only openai key configured — fallback should skip groq silently.
fake_config.set("openai_api_key", "sk-test")
calls: List[str] = []
def fake_post(url, headers=None, files=None, data=None, timeout=None):
calls.append(url)
return FakeResponse(200, "via-openai")
monkeypatch.setattr(tr.requests, "post", fake_post)
text = tr._transcribe_with_fallback(chunk_file, ["groq", "openai"], fake_config)
assert text == "via-openai"
assert calls == [tr.PROVIDERS["openai"]["endpoint"]]
def test_all_fail_raises_with_last_error(self, monkeypatch, fake_config, chunk_file):
fake_config.set("groq_api_key", "gsk_test")
fake_config.set("openai_api_key", "sk-test")
monkeypatch.setattr(
tr.requests,
"post",
lambda *a, **k: FakeResponse(500, "boom"),
)
with pytest.raises(tr.TranscribeError, match="all providers failed"):
tr._transcribe_with_fallback(chunk_file, ["groq", "openai"], fake_config)
# --- transcribe (orchestrator) ---------------------------------------- #
class TestOrchestrator:
def test_local_file_skips_yt_dlp(self, monkeypatch, fake_config, tmp_path, chunk_file):
fake_config.set("groq_api_key", "gsk_test")
def boom_download(*a, **k):
raise AssertionError("yt-dlp must not be called for local files")
# Stub heavy external steps to no-ops that keep file paths valid.
compressed = tmp_path / "compressed.m4a"
compressed.write_bytes(b"x" * 1024)
def fake_compress(src, out_dir):
return compressed
monkeypatch.setattr(tr, "download_audio", boom_download)
monkeypatch.setattr(tr, "compress_audio", fake_compress)
monkeypatch.setattr(
tr.requests,
"post",
lambda *a, **k: FakeResponse(200, "transcript text"),
)
text = tr.transcribe(
str(chunk_file),
out_dir=tmp_path / "work",
config=fake_config,
)
assert text == "transcript text"
def test_chunks_concatenated_with_newlines(
self, monkeypatch, fake_config, tmp_path, chunk_file
):
fake_config.set("groq_api_key", "gsk_test")
# Force the "needs chunking" path by writing a file above the size limit.
big = tmp_path / "compressed.m4a"
big.write_bytes(b"x" * (tr.SIZE_LIMIT_BYTES + 1))
monkeypatch.setattr(tr, "compress_audio", lambda src, out_dir: big)
c1 = tmp_path / "chunk_001.m4a"
c2 = tmp_path / "chunk_002.m4a"
c1.write_bytes(b"a")
c2.write_bytes(b"b")
monkeypatch.setattr(tr, "chunk_audio", lambda src, out_dir: [c1, c2])
responses = iter(["part one ", "part two "])
monkeypatch.setattr(
tr.requests,
"post",
lambda *a, **k: FakeResponse(200, next(responses)),
)
text = tr.transcribe(
str(chunk_file),
out_dir=tmp_path / "work",
config=fake_config,
)
assert text == "part one\npart two"
def test_no_provider_configured_fails_fast(self, fake_config, chunk_file):
with pytest.raises(tr.NoProviderConfigured):
tr.transcribe(str(chunk_file), config=fake_config)
def test_invalid_provider_string(self, fake_config, chunk_file):
with pytest.raises(tr.TranscribeError, match="unknown provider"):
tr.transcribe(str(chunk_file), provider="azure", config=fake_config)
# --- YouTubeChannel integration --------------------------------------- #
class TestYouTubeChannelTranscribe:
def test_delegates_to_transcribe(self, monkeypatch, fake_config):
from agent_reach.channels.youtube import YouTubeChannel
captured = {}
def fake_transcribe(source, *, provider="auto", out_dir=None, config=None):
captured["source"] = source
captured["provider"] = provider
captured["config"] = config
return "delegated text"
monkeypatch.setattr(tr, "transcribe", fake_transcribe)
out = YouTubeChannel().transcribe(
"https://youtu.be/abc", provider="groq", config=fake_config
)
assert out == "delegated text"
assert captured["source"] == "https://youtu.be/abc"
assert captured["provider"] == "groq"
assert captured["config"] is fake_config
# --- Config feature requirement --------------------------------------- #
class TestConfigOpenAIWhisper:
def test_openai_whisper_feature_registered(self, fake_config):
assert "openai_whisper" in Config.FEATURE_REQUIREMENTS
assert Config.FEATURE_REQUIREMENTS["openai_whisper"] == ["openai_api_key"]
assert not fake_config.is_configured("openai_whisper")
fake_config.set("openai_api_key", "sk-test")
assert fake_config.is_configured("openai_whisper")
-46
View File
@@ -26,7 +26,6 @@ def test_check_twitter_cli_found_and_auth_ok():
assert status == "ok" assert status == "ok"
assert "twitter-cli" in message assert "twitter-cli" in message
assert "完整可用" in message assert "完整可用" in message
assert channel.active_backend == "twitter-cli"
def test_check_twitter_cli_found_auth_missing(): def test_check_twitter_cli_found_auth_missing():
@@ -42,8 +41,6 @@ def test_check_twitter_cli_found_auth_missing():
status, message = channel.check() status, message = channel.check()
assert status == "warn" assert status == "warn"
assert "未认证" in message assert "未认证" in message
# 未认证是业务态:工具进程活着,后端仍可用
assert channel.active_backend == "twitter-cli"
# --- bird CLI fallback tests --- # --- bird CLI fallback tests ---
@@ -62,7 +59,6 @@ def test_check_bird_fallback_auth_ok():
status, message = channel.check() status, message = channel.check()
assert status == "ok" assert status == "ok"
assert "bird" in message assert "bird" in message
assert channel.active_backend == "bird CLI (legacy)"
def test_check_bird_fallback_auth_missing(): def test_check_bird_fallback_auth_missing():
@@ -90,7 +86,6 @@ def test_check_nothing_installed():
status, message = channel.check() status, message = channel.check()
assert status == "warn" assert status == "warn"
assert "twitter-cli" in message assert "twitter-cli" in message
assert channel.active_backend is None
# --- twitter-cli preferred over bird --- # --- twitter-cli preferred over bird ---
@@ -111,44 +106,3 @@ def test_twitter_cli_preferred_over_bird():
status, message = channel.check() status, message = channel.check()
assert status == "ok" assert status == "ok"
assert "twitter-cli" in message assert "twitter-cli" in message
assert channel.active_backend == "twitter-cli"
# --- broken install (stale venv shim) ---
def test_check_twitter_cli_broken_reports_error_with_reinstall_hint():
"""which 命中但 exec 抛 FileNotFoundErrorvenv 断链)→ error + 重装处方。"""
channel = TwitterChannel()
with patch(
"shutil.which",
side_effect=lambda name: "/usr/local/bin/twitter" if name == "twitter" else None,
), patch("subprocess.run", side_effect=FileNotFoundError("/usr/local/bin/twitter")):
status, message = channel.check()
assert status == "error"
assert "无法执行" in message
assert "uv tool install --force twitter-cli" in message
assert "pipx reinstall twitter-cli" in message
assert channel.active_backend is None
def test_check_twitter_cli_broken_falls_back_to_bird():
"""twitter-cli 断链但 bird 健康 → 回退到 bird,后端正确归属。"""
channel = TwitterChannel()
def which_side_effect(name):
if name in ("twitter", "bird"):
return f"/usr/local/bin/{name}"
return None
def run_side_effect(cmd, **kwargs):
if "twitter" in cmd[0]:
raise FileNotFoundError(cmd[0])
return _cp(stdout="Authenticated as @user\n", returncode=0)
with patch("shutil.which", side_effect=which_side_effect), patch(
"subprocess.run", side_effect=run_side_effect
):
status, message = channel.check()
assert status == "ok"
assert "bird" in message
assert channel.active_backend == "bird CLI (legacy)"
Generated
+1640
View File
File diff suppressed because it is too large Load Diff