osint-dashboard/docs/news.md

140 lines
5.6 KiB
Markdown
Raw Permalink Normal View History

# News pipeline — scraper + summarizer
The OSINT dashboard ingests ~257 global news RSS sources hourly and produces
LLM master summaries. Both services were vendored from the upstream
`~/Projects/newsPipeline` project and re-integrated here to replace the old
k8s CronJob choreography with in-compose scheduling against the EXISTING
osint-db — **no second Postgres**.
## Architecture
```
257 RSS feeds (news/scraper/urls.txt)
news-scraper (Scrapy, hourly :00) ──► articles table (osint-db)
│ │
│ ▼
news-summarizer (Gemini map-reduce, hourly :05) ──► article_summaries table
GET /api/news · GET /api/news/summaries
```
| Component | Image | Container | Scheduling |
|---|---|---|---|
| Scraper | `localhost/osint-news-scraper` | `osint-news-scraper` | wall-clock loop, minute `NEWS_SCRAPE_MINUTE` (default :00) |
| Summarizer | `localhost/osint-news-summarizer` | `osint-news-summarizer` | wall-clock loop, minute `NEWS_SUMMARIZE_MINUTE` (default :05) |
Both services live under the `ingest` compose profile (same as the ingester
and camera-scraper): `docker compose --profile ingest up -d`.
## Data flow
1. **Scraper**`news/scraper/run_news_scraper.py` runs
`scrapy crawl articles` (spider `news/scraper/newsScraper/spiders/news_spider.py`)
at the top of each hour. The spider reads the RSS feed URLs from `urls.txt`,
follows each `<item>` link, extracts the main article body, and the
`PostgresPipeline` writes to `articles` with URL-based dedup
(`ON CONFLICT (url) DO NOTHING`).
2. **Summarizer**`news/summerizer/run_news_summarizer.py` runs
`summarizer.py` at :05 past each hour. It reads articles from the last
`SUMMARY_WINDOW_HOURS`, map-reduces them through Gemini
(`SUMMARY_MODEL`, default `gemini-2.0-flash`), and inserts one master
summary into `article_summaries`.
Scheduling is done with small in-compose wall-clock loops (not host cron): each
loop runs once on boot (`*_RUN_ON_START=1`, seeds data fast) then sleeps until
the next scheduled minute. The loop is serial, so a run that overruns its slot
simply shifts to the next boundary — two crawls/summaries never overlap.
The `articles` and `article_summaries` tables are created by the idempotent
alembic migration `003_news` (also created by the scraper's own
`CREATE TABLE IF NOT EXISTS`, so container startup order doesn't matter).
## Endpoints
### GET /api/news — recent articles
| Query param | Meaning | Default |
|---|---|---|
| `domain` | filter by source domain (e.g. `www.reuters.com`) | none |
| `since` | only articles captured at/after this UTC instant (ISO-8601) | none |
| `limit` | max rows | `50` (max `500`) |
| `offset` | pagination offset | `0` |
```json
[
{
"id": 1,
"title": "…",
"url": "https://…",
"content": "full extracted article text…",
"domain": "www.reuters.com",
"timestamp": "2026-08-24T18:10:00Z"
}
]
```
### GET /api/news/summaries — master LLM summaries
| Query param | Meaning | Default |
|---|---|---|
| `since` | only summaries generated at/after this UTC instant | none |
| `limit` | max rows | `20` (max `100`) |
| `offset` | pagination offset | `0` |
```json
[
{
"id": 1,
"summary_text": "master LLM summary (markdown)…",
"batch_timestamp": "2026-08-24T18:10:00Z"
}
]
```
## Configuration (all via env / `.env`)
| Var | Default | Notes |
|---|---|---|
| `GEMINI_API_KEY` | *(blank)* | **Required for summaries.** Unset = summarizer logs and idles (never crashes). |
| `SUMMARY_MODEL` | `gemini-2.0-flash` | Gemini model id. |
| `NEWS_BATCH_SIZE` | `50` | Articles per map-phase batch. |
| `SUMMARY_WINDOW_HOURS` | `1` | How far back the summarizer looks for new articles. |
| `INCLUDE_FUTURES` | `0` | Legacy futures-prices coupling (upstream pipeline). OFF for OSINT; set `1` + install `yfinance` to enable. |
| `NEWS_SCRAPE_MINUTE` | `0` | Wall-clock minute the scraper fires. |
| `NEWS_SUMMARIZE_MINUTE` | `5` | Wall-clock minute the summarizer fires. |
| `NEWS_SCRAPE_RUN_ON_START` | `1` | Run one scrape immediately on container start. |
| `NEWS_SUMMARIZE_RUN_ON_START` | `1` | Run one summarize immediately on container start. |
| `NEWS_LOG_LEVEL` | `INFO` | Scrapy log level. |
| `TELEGRAM_TOKEN` / `TELEGRAM_CHAT_ID` | *(blank)* | Reserved for the (out-of-scope) Telegram delivery bot. |
DB_* for both services is mapped to the shared osint-db credentials
(`DB_HOST=db`, same `DB_USER/DB_PASSWORD/DB_NAME` as the rest of the stack).
## Prompts
Both prompts are env-overridable — the default `MAP_PROMPT` is OSINT-neutral
(facts, locations, entities, category, OSINT signal per article) and the default
`SUMMARY_PROMPT` produces a concise executive summary of the most impactful
items (with a "no qualifying events" escape hatch). Upstream's futures/markets
prompt language is gated behind `INCLUDE_FUTURES=1`.
## Tests
`tests/test_api_news.py` — DB-backed API contract tests (auto-skip without a
reachable test database, same as the FIRMS tests):
```bash
DB_HOST=... DB_PORT=... DB_USER=osint DB_PASSWORD=... DB_NAME=osint_data \
pytest tests/test_api_news.py -v
```
## Live verification
End-to-end (real crawl → DB → API) is verified after deploy on the Pi: check
`docker compose --profile ingest logs -f news-scraper news-summarizer`, then
`curl -s localhost:8000/api/news | head`. Summaries additionally require
`GEMINI_API_KEY` to be set in `.env` on the Pi.