Fetch articles from local WeWe-RSS service, read full content via Playwright, clean and output structured data...
This skill retrieves articles from a local RSS service (WeWe-RSS) and outputs clean, structured content ready for downstream processing (e.g., video script generation).
Follow these steps EXACTLY when executing this skill:
Analyze user input to determine mode:
Use the provided python script to fetch RSS articles:
python3 app/agentic/skills/rss-article-retriever/scripts/fetch_rss.py --limit {limit} --title_include "{keyword}"
Use the provided python script which uses Playwright:
python3 app/agentic/skills/rss-article-retriever/scripts/fetch_content.py "{article_url}"
If these scripts are available as tools, use the tool directly. Otherwise execute via shell.
Apply cleaning rules from CLEANING_RULES.md:
Look for common legal article headings:
Wrap output in <skill-output> markers with proper JSON format.
MANDATORY: After outputting the RSS article result, check the original user request:
If the user's request contains ANY of these phrases:
You MUST IMMEDIATELY call the video-script-generator skill with the article content.
DO NOT ask the user "็ฐๅจ้่ฆ็จ่ฟ็ฏๆ็ซ ๆฅๅถไฝ่ง้ข่ๆฌๅ๏ผ" or similar questions.
DO NOT stop after outputting the article.
Execute this command:
Skill(video-script-generator)
With the article's plain_text content as input.
Example of correct behavior:
User: "่ฏทๆพไธ็ฏๅ
ณไบ'ๆฐ้ดๅ่ดทๅๅ'็ๆ็ซ ๅถไฝ่ง้ข่ๆฌ"
Agent actions:
1. Search RSS for articles about "ๆฐ้ดๅ่ดทๅๅ"
2. Get full content using Python Playwright
3. Output <skill-output type="rss-article">...</skill-output>
4. IMMEDIATELY call: Skill(video-script-generator) with the article content
5. Output <skill-output type="video-script">...</skill-output>
โ DO NOT use Python requests library - WeChat pages are dynamically rendered with JavaScript
โ DO NOT use import requests or from bs4 import BeautifulSoup - These won't work for dynamic content
โ DO NOT say "Playwright is unavailable" or "tool usage limit reached" - The Python Playwright method above works
โ DO NOT give up and ask user to provide content - Always use the Python Playwright method first
โ DO NOT use curl to parse WeChat article content - WeChat articles are JavaScript-rendered
โ
USE Python Playwright via Bash python3 -c - This is the ONLY reliable method for WeChat articles
โ
USE asyncio and playwright.async_api - These handle async browser operations
โ
USE page.evaluate() to extract content - This gets the rendered text content
โ
USE wait_until='networkidle' - This ensures the page is fully loaded before extraction
http://localhost:4000/feeds/all.atom<skill-output> boundary markers| Mode | Trigger Keywords | Action | Output |
|---|---|---|---|
list |
"ๅ่กจ", "ๆๅชไบ", "ๆ่ฟๆ็ซ ", "list" | Return article list (no content) | article_summary[] |
search |
"ๅ ณไบ", "ๅ ๅซ", "ๆ็ดข" + keyword | Filter by title keyword | article_summary[] |
get |
"ๆๆฐ", "็ฌฌN็ฏ", "ๅ ๅฎน", "ๅฎๆด", "get" | Fetch full content via Playwright | article_full |
Endpoint: http://localhost:4000/feeds/all.atom
Parameters:
| Parameter | Type | Description |
|---|---|---|
limit |
number | Number of articles to return (default: 10) |
page |
number | Page number for pagination |
title_include |
string | Filter: title must include this keyword |
title_exclude |
string | Filter: title must not include this keyword |
Article URL Pattern: https://mp.weixin.qq.com/s/{article.id}
Example Calls:
# Get latest 5 articles
curl "http://localhost:4000/feeds/all.atom?limit=5"
# Search articles containing "ๆฐ้ดๅ่ดท"
curl "http://localhost:4000/feeds/all.atom?title_include=ๆฐ้ดๅ่ดท"
When fetching full article content, use one of these methods:
python3 -c "
import asyncio
from playwright.async_api import async_playwright
async def main():
async with async_playwright() as p:
browser = await p.chromium.launch(headless=True)
page = await browser.new_page()
await page.goto('https://mp.weixin.qq.com/s/{article_id}', wait_until='networkidle', timeout=30000)
try:
await page.wait_for_selector('#js_content', timeout=15000)
except:
pass
content = await page.evaluate('() => { const el = document.getElementById(\"js_content\"); return el ? el.innerText : \"\"; }')
print(content)
await browser.close()
asyncio.run(main())
" 2>&1
# For WeChat articles, try direct fetch first
curl -s "https://mp.weixin.qq.com/s/{article_id}" -H "User-Agent: Mozilla/5.0" | \
grep -A 1000 'id="js_content"' | grep -B 1000 'id="sg_' | \
sed 's/<[^>]*>//g' | sed 's/ / /g' | sed 's/</</g' | sed 's/>/>/g' | tr -s '\n'
If previous methods fail, check if content was saved and read from temp file.
Content Extraction Steps:
#js_content divSee CLEANING_RULES.md for detailed rules.
Remove these elements:
#js_profile_qrcode, .profile_container)#js_pc_qr_code, .qr_code_pc).rich_media_tool)#js_sponsor_ad_area)#js_profile_article)Preserve these elements:
p, section with text)strong, b, h1-h6)ul, ol, li)Section Recognition (Legal Articles): Common headings to identify sections:
All structured output MUST be wrapped in boundary markers:
<skill-output type="rss-article" schema-version="1.0.0">
{JSON conforming to schema/output-schema-v1.json}
</skill-output>
Marker Attributes:
| Attribute | Value | Description |
|---|---|---|
type |
rss-article |
Fixed value identifying skill type |
schema-version |
1.0.0 |
Schema version for compatibility |
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ Skill Complete Output โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค
โ [Optional] Explanatory text (should be ignored) โ
โ โ
โ <skill-output type="rss-article" schema-version="1.0.0">
โ { โ
โ "version": "1.0.0", โ
โ "operation": "get", โ
โ "result": { ... } โ Structured result โ
โ } โ
โ </skill-output> โ
โ โ
โ [Optional] Additional notes (should be ignored) โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
Main Agent Parsing Logic:
<skill-output ...> and </skill-output>list and search Operations{
"version": "1.0.0",
"operation": "list",
"result": {
"type": "summary",
"total_count": 5,
"articles": [
{
"id": "article-id-123",
"title": "Article Title",
"author": "ๅ
ฌไผๅทๅ็งฐ",
"published_at": "2026-01-20T10:30:00Z",
"url": "https://mp.weixin.qq.com/s/article-id-123",
"summary": "RSS ๆ่ฆๅ
ๅฎน..."
}
],
"query": {
"limit": 5,
"page": 1,
"title_include": null,
"title_exclude": null
}
}
}
get Operation{
"version": "1.0.0",
"operation": "get",
"result": {
"type": "full",
"article": {
"id": "article-id-123",
"title": "Article Title",
"author": "ๅ
ฌไผๅทๅ็งฐ",
"published_at": "2026-01-20T10:30:00Z",
"url": "https://mp.weixin.qq.com/s/article-id-123",
"content": {
"plain_text": "ๆธ
ๆดๅ็็บฏๆๆฌ๏ผๅฏ็ดๆฅไผ ้็ป video-script-generator",
"sections": [
{ "heading": "ๆกๆ
็ฎไป", "text": "..." },
{ "heading": "ๆณ้ขๅฎก็", "text": "..." }
],
"source": "playwright"
},
"word_count": 1245,
"cleaning_notes": ["Removed author profile", "Removed QR codes"]
}
}
}
| Mode | Trigger | Output Content | Use Case |
|---|---|---|---|
| Default | No parameter | Boundary markers + JSON + optional notes | SaaS integration, main agent calls |
| Raw | --raw |
Pure JSON (no markers, no notes) | Direct API calls, testing |
Before generating output, complete these checks:
Read Schema File
schema/output-schema-v1.jsonReference Examples
references/EXAMPLES.mdFormat Requirements
version must use semantic versioning: "1.0.0"published_at must use ISO 8601 formatword_count must be calculated from cleaned content{
"version": "1.0.0",
"operation": "list",
"error": {
"code": "RSS_CONNECTION_ERROR",
"message": "Cannot connect to RSS service at localhost:4000",
"suggestion": "Ensure WeWe-RSS service is running"
}
}
{
"version": "1.0.0",
"operation": "get",
"result": {
"type": "full",
"article": {
"id": "...",
"content": {
"plain_text": "RSS ๆ่ฆไฝไธบๅ้ๅ
ๅฎน...",
"sections": [],
"source": "rss_fallback"
},
"cleaning_notes": ["Playwright timeout - using RSS summary as fallback"]
}
}
}
| Code | Description |
|---|---|
RSS_CONNECTION_ERROR |
Cannot connect to RSS service |
RSS_PARSE_ERROR |
Failed to parse RSS feed |
ARTICLE_NOT_FOUND |
Article ID not found in feed |
PLAYWRIGHT_TIMEOUT |
Page load timeout (fallback to RSS summary) |
CONTENT_EMPTY |
No content extracted from page |
When user requests "็จRSSๆ็ซ ๅถไฝ่ง้ข่ๆฌ":
get operationresult.article.content.plain_textExample Chain:
User: "่ฏท็จๆๆฐไธ็ฏRSSๆ็ซ ๆฅๅถไฝ่ง้ข่ๆฌ"
Step 1: rss-article-retriever (get latest)
โ Outputs: <skill-output type="rss-article">...</skill-output>
Step 2: Extract plain_text from result
Step 3: video-script-generator (with plain_text as input)
โ Outputs: <skill-output type="video-script">...</skill-output>
User input: "่ทๅๆ่ฟ5็ฏๆ็ซ ๅ่กจ"
Output:
<skill-output type="rss-article" schema-version="1.0.0">
{
"version": "1.0.0",
"operation": "list",
"result": {
"type": "summary",
"total_count": 5,
"articles": [
{
"id": "abc123",
"title": "ๆฐ้ดๅ่ดท็บ ็บทๆกไพๅๆ",
"author": "ๆณๅพๅ
ฌไผๅท",
"published_at": "2026-01-20T10:30:00Z",
"url": "https://mp.weixin.qq.com/s/abc123",
"summary": "ๆฌๆกๆถๅๆฐ้ดๅ่ดท..."
},
...
],
"query": { "limit": 5, "page": 1 }
}
}
</skill-output>
User input: "ๆพไธ็ฏๅ ณไบ'ๆฐ้ดๅ่ดท'็ๆ็ซ "
Output:
<skill-output type="rss-article" schema-version="1.0.0">
{
"version": "1.0.0",
"operation": "search",
"result": {
"type": "summary",
"total_count": 2,
"articles": [...],
"query": { "title_include": "ๆฐ้ดๅ่ดท" }
}
}
</skill-output>
User input: "่ทๅๆๆฐไธ็ฏๆ็ซ ็ๅฎๆดๅ ๅฎน"
Output:
<skill-output type="rss-article" schema-version="1.0.0">
{
"version": "1.0.0",
"operation": "get",
"result": {
"type": "full",
"article": {
"id": "abc123",
"title": "ๆฐ้ดๅ่ดท็บ ็บทๆกไพๅๆ",
"author": "ๆณๅพๅ
ฌไผๅท",
"published_at": "2026-01-20T10:30:00Z",
"url": "https://mp.weixin.qq.com/s/abc123",
"content": {
"plain_text": "ๆกๆ
็ฎไป\n\nๅผ ๆไธๆๆ็ณปๆๅๅ
ณ็ณป...",
"sections": [
{ "heading": "ๆกๆ
็ฎไป", "text": "ๅผ ๆไธๆๆ็ณปๆๅๅ
ณ็ณป..." },
{ "heading": "ๆณ้ขๅฎก็", "text": "ๆณ้ข็ปๅฎก็่ฎคไธบ..." },
{ "heading": "ๆณ้ขๅคๅณ", "text": "ๅคๅณ่ขซๅๆๆ..." }
],
"source": "playwright"
},
"word_count": 1245,
"cleaning_notes": ["Removed author profile", "Removed QR codes", "Removed subscription prompt"]
}
}
}
</skill-output>