Quatro fases do roadmap do refactor de frontend em uma unica entrega. Fase 1 — NSFW gate visual - get_media_files / get_total_media_count / get_recent_downloads / get_media_by_authors retornam nsfw (+ flair, source_type, created_utc ja persistidos). - GET /api/media aceita ?nsfw=all|hide|only com regex validation. - Dropdown "NSFW: tudo / ocultar / apenas" na barra de filtros da galeria. - appendToGallery aplica .nsfw-thumb (blur 22px) quando item.nsfw. - Badge "NSFW" vermelho no canto do thumbnail. - Toggle 👁/🙈 no header (Alpine.store.prefs.showNsfw) revela o blur globalmente; clase .show-nsfw no body. Fase 2 — Tag taxonomy estilo Stash - Schema novo: tabelas tags (id, name, category, is_nsfw, description) e post_tags (post_id, tag_id, source). Indexes em ambas direcoes. - auto_tag_post() popula 4 categorias automaticas a cada add_post: subreddit -> 'source', author quando source_type=user -> 'performer', flair -> 'genre', nsfw=1 -> 'meta'. Idempotente; preserva tags criadas pelo usuario (source='user'). - Novo router src/web/routers/tags.py: - GET /api/tags (com ?category) - GET /api/posts/{id}/tags - POST /api/posts/{id}/tags (rate-limited) - DELETE /api/posts/{id}/tags/{tag_id} - POST /api/tags/backfill — corre auto_tag_post em todos os posts existentes (idempotente). - main.py: process_post agora chama db.auto_tag_post(record). Fase 3 — PIN lock + session cookie - Novo src/web/session.py: cookie 'rmc_unlock' assinado por HMAC-SHA256 com chave de 32 bytes gerada por boot (in-memory; restart invalida sessoes). Format: '{issued_unix}.{hex_sig}'. - Env vars novas: RMC_PIN (4-6 digitos, se ausente desliga o lock inteiro), RMC_PIN_TIMEOUT (segundos, default 600). - PinLockMiddleware: gateia tudo exceto /health, /static, /unlock, /favicon. GET sem cookie redirect 303 /unlock; non-GET sem cookie retornam 401. - Endpoints: GET /unlock renderiza tela com input password OTP-friendly; POST /unlock valida via secrets.compare_digest e seta cookie httponly samesite=lax (secure auto quando https). POST /lock invalida cookie. - Template unlock.html (sem auth/cookie), CSS .unlock-wrap/.unlock-card. - Header ganha botao "🔒 Lock" quando pin_enabled (passado via context). - Nova dep python-multipart (necessaria pra Form() do FastAPI). Fase 4 — Galeria enriquecida + atalhos - Badges Reddit-specific no card: NSFW, source_type icon (👤/📂), tempo relativo created_utc (agora/Xm/Xh/Xd/Xmes/Xa), chip de flair com estilo proprio. - Keyboard shortcuts no modal: j/→ proximo, k/← anterior, f favoritar, b blacklist autor, Esc fecha. Inativo quando focus em input. - Modo discreto turbinado: grids compactas (.gallery-grid e .authors-grid em minmax(110px)), thumbs menores em recent, info reduzida, flair-badge escondida. Discreet-banner discreto no topo. - Auto-discreet: 60s sem mouse/keyboard/click/scroll → ativa discreet automaticamente. Privacidade hands-off. Plus - escapeHtml() helper para titulos e nomes nos templates JS. - CSS .tag-chip preparado para Fase 5 (chips de tag no card visual); estilizado por categoria (performer/source/genre/meta). Bump 1.2.1 → 1.3.0 (minor — 4 features novas, schema novo, nova env var, nova dep). Backwards-compat preservado: clientes JSON antigos continuam recebendo JSON; APIs nao quebraram contratos. Verificado: 151 testes verdes, ruff/mypy/format OK, docker build OK, smoke manual cobre PIN lock (redirect, 401, cookie issue, bypass de health/static), backfill (11 posts tagged), nsfw filter (hide retorna so nsfw=0), keyboard shortcuts. |
||
|---|---|---|
| .github/workflows | ||
| src | ||
| tests | ||
| .dockerignore | ||
| .gitignore | ||
| .pre-commit-config.yaml | ||
| config.yaml.example | ||
| docker-compose.synology.yml | ||
| docker-compose.yml | ||
| Dockerfile | ||
| pyproject.toml | ||
| README.md | ||
| reddit-collector-web.service | ||
| run.sh | ||
| run_collector.sh | ||
| run_web.sh | ||
Reddit Media Collector
A powerful, self-hosted media collector for Reddit that automatically downloads images, videos, and GIFs from your favourite subreddits and users. Features a built-in web interface for management and seamless integration with Immich for photo organisation.
Features
- Multi-source Collection - Collect from subreddits and user profiles
- Smart Deduplication - MD5 hash-based detection prevents duplicate downloads
- Gallery Support - Automatically handles Reddit galleries with multiple images
- Multiple Extractors - Built-in support for Reddit, Imgur, Gfycat, and Redgifs
- Immich Integration - Generates JSON sidecar files with metadata for seamless import
- Web Dashboard - Modern web interface for configuration and monitoring
- Blacklist System - Filter out unwanted authors, subreddits, keywords, and domains
- Favourites System - Mark and filter your favourite posts
- Video Thumbnails - Auto-generated thumbnails for video preview in gallery
- No API Keys Required - Uses Reddit's public JSON endpoints
- Docker Support - Easy deployment with Docker Compose
- Scheduled Collection - Cron-ready for automated periodic collection
Quick Start
Prerequisites
- Python 3.11 or higher
- FFmpeg (optional, for video thumbnails)
- yt-dlp (optional, for Gfycat/Redgifs support)
Installation
-
Clone the repository
git clone https://github.com/richardnixondev/reddit-media-collector.git cd reddit-media-collector -
Create virtual environment
python3 -m venv venv source venv/bin/activate # Linux/macOS # or .\venv\Scripts\activate # Windows -
Install the package (editable, with dev tooling)
pip install -e ".[dev]" -
Configure
cp config.yaml.example config.yaml # Edit config.yaml with your preferences -
Run the collector
python -m src.main
Development
pip install -e ".[dev]"
pre-commit install
# Tests + coverage
pytest -v
pytest --cov=src --cov-report=term-missing
# Lint, format, type-check
ruff check src/ tests/
ruff format src/ tests/
mypy src/
Configuration
Create a config.yaml file based on the example:
targets:
subreddits:
- name: "earthporn"
limit: 100 # Posts per request (max 100)
sort: "top" # hot, new, top, rising
time_filter: "week" # For "top" sort: hour, day, week, month, year, all
- name: "pics"
limit: 50
sort: "new"
users:
- name: "username"
limit: 100
download:
output_dir: "./downloads"
media_types:
- "image"
- "video"
- "gif"
min_score: 10 # Minimum upvotes required
skip_nsfw: false # Skip NSFW content
max_file_size_mb: 200 # Maximum file size
flat_structure: true # All files in a single folder
generate_sidecar: true # Generate .json for Immich
videos_only_from_favorites: false # Only download videos from favourited users
rate_limit:
requests_per_minute: 20
download_delay_seconds: 2
logging:
level: "INFO"
file: "collector.log"
blacklist:
authors: [] # Usernames to ignore
subreddits: [] # Subreddits to ignore
title_keywords: []
domains: []
Web Interface
The collector includes a FastAPI-powered web dashboard for easy management.
Starting the Web Server
# Using the run script
./run_web.sh
# Or directly with uvicorn
uvicorn src.web.app:app --host 0.0.0.0 --port 8000
Access the dashboard at http://localhost:8000
Web Features
- Dashboard - View collection statistics and manage targets
- Gallery - Browse downloaded media with filtering, infinite scroll, and favourites
- Authors - Browse content grouped by author, with per-author modal
- Settings - Configure download options, blacklist, and scheduler
- Scheduler - Configure recurring collection runs (replaces external cron when running as a service)
- Collector Control - Trigger collection runs manually or per target
Security & Limits
HTTP Basic Auth (optional)
Set both RMC_AUTH_USER and RMC_AUTH_PASS to require Basic credentials on every endpoint (including /). With either variable unset, the API stays public — appropriate for trusted local/intranet deployments.
RMC_AUTH_USER=alice RMC_AUTH_PASS=s3cret uvicorn src.web.app:app
Rate limiting (per IP, per endpoint)
Heavy endpoints are throttled in-process to prevent runaway clients:
| Endpoint | Limit |
|---|---|
GET /api/stats, /api/stats/enhanced |
30 / minute |
POST /api/collector/run |
3 / minute |
POST /api/media/cleanup-blacklist |
5 / minute |
POST /api/media/cleanup-by-type |
5 / minute |
Limits are per (client IP, route path) and reset on a sliding window. For multi-worker deployments, swap the in-memory buckets for a shared backend.
Gallery performance
- The gallery keeps at most 500 items in memory/DOM at once; older items are evicted as you scroll. Use filters or sort to reach older media.
- Status polling pauses when the tab is hidden.
- Filter dropdowns are debounced (200 ms) so rapid changes only fire one request.
- Bulk DELETE runs 5 requests in parallel.
Running as a Service (systemd)
# Copy the service file
sudo cp reddit-collector-web.service /etc/systemd/system/
# Reload systemd and enable service
sudo systemctl daemon-reload
sudo systemctl enable reddit-collector-web
sudo systemctl start reddit-collector-web
# Check status
sudo systemctl status reddit-collector-web
# View logs
sudo journalctl -u reddit-collector-web -f
Docker
Using Docker Compose
# Build and run
docker-compose up -d
# View logs
docker-compose logs -f
# Stop
docker-compose down
Docker Configuration
The docker-compose.yml mounts the following volumes:
./config.yaml→/app/config.yaml(must be writable — the UI mutates it)./downloads→/app/downloads(media files)./data→/app/data(SQLite DB, scheduler DB/config — survives upgrades)
Relevant environment variables (all optional, with sensible defaults inside the image):
| Variable | Default | Purpose |
|---|---|---|
RMC_DOWNLOAD_DIR |
/app/downloads |
Where media files are written |
RMC_DB_PATH |
/app/data/media.db |
Main SQLite database |
RMC_SCHEDULER_DB |
/app/scheduler.db |
APScheduler jobstore — set to /app/data/scheduler.db in containers so history survives restarts |
RMC_SCHEDULER_CONFIG |
/app/scheduler_config.yaml |
Scheduler interval/cron config — same advice as above |
RMC_CONFIG_PATH |
/app/config.yaml |
YAML with subreddits/users/blacklist |
RMC_TIMEZONE |
UTC |
Timezone used by the scheduler |
RMC_AUTH_USER / RMC_AUTH_PASS |
unset | Enable HTTP Basic Auth on every route when both are set |
Synology DSM Deployment
Tested on DSM 7.2 with Container Manager on x86_64 Plus models. The published image lives at
ghcr.io/richardnixondev/reddit-media-collector:latest.
-
Prepare the host directories (one-time, via SSH or File Station):
sudo mkdir -p /volume1/docker/reddit-media-collector/{downloads,data} sudo cp /path/to/your/config.yaml /volume1/docker/reddit-media-collector/config.yaml -
Create the project in Container Manager:
- Open Container Manager → Project → Create.
- Project name:
reddit-media-collector. - Path:
/volume1/docker/reddit-media-collector. - Source: Create docker-compose.yml, paste the contents of
docker-compose.synology.yml(adjustRMC_TIMEZONEto your zone). - Next → Build — DSM does
docker compose pull && up -d.
-
Access the UI:
http://<nas-ip>:8000. Downloaded media appears in/volume1/docker/reddit-media-collector/downloads/, which you can point Synology Photos or an Immich instance at. -
HTTPS via DSM Reverse Proxy (optional, recommended if exposed): Control Panel → Login Portal → Advanced → Reverse Proxy → Create. Source:
reddit.yourdomain.com(HTTPS:443). Destination:localhost:8000(HTTP). Attach a Let's Encrypt cert. When exposing publicly, uncommentRMC_AUTH_USER/RMC_AUTH_PASSin the compose file. -
Updating: Container Manager → Project → reddit-media-collector → Action → Build re-pulls
latest. Thedata/volume keeps the database, scheduler history and config.
Permissions note: if the container can't write to the bind mounts, check the owner
of /volume1/docker/reddit-media-collector/ with ls -ln. Either chown it to a user
the container can write as, or set user: "UID:GID" in the compose (commented out at the
bottom of docker-compose.synology.yml).
File Naming Convention
Downloaded files follow a descriptive naming pattern:
{subreddit}_{author}_{YYYYMMDD}_{HHmmss}_{post_id}[_{gallery_index}].{ext}
Examples:
earthporn_photographer123_20260118_143052_abc123.jpg
pics_user456_20260118_091523_xyz789_1.jpg (gallery image 1)
pics_user456_20260118_091523_xyz789_2.jpg (gallery image 2)
Immich Integration
The collector generates JSON sidecar files compatible with Immich:
{
"dateTimeOriginal": "2026-01-18T14:30:52+00:00",
"description": "Post title from Reddit",
"albums": ["r/earthporn"],
"tags": ["reddit", "earthporn", "image"],
"rating": 4,
"people": ["photographer123"],
"externalUrl": "https://reddit.com/r/earthporn/comments/abc123/title"
}
Rating System
Ratings are automatically assigned based on post score:
| Score | Rating |
|---|---|
| 0-9 | 1 star |
| 10-49 | 2 stars |
| 50-199 | 3 stars |
| 200-999 | 4 stars |
| 1000+ | 5 stars |
Importing to Immich
Point Immich to your downloads folder as an external library, or use the Immich CLI:
immich upload --album "Reddit Collection" ./downloads/
Scheduled Collection
Using Cron
# Edit crontab
crontab -e
# Add entry (runs every 6 hours)
0 */6 * * * /path/to/reddit-media-collector/run_collector.sh
Example run_collector.sh
#!/bin/bash
cd /path/to/reddit-media-collector
source venv/bin/activate
export PATH="$HOME/.local/bin:$PATH" # For yt-dlp
timeout 4h python -m src.main >> cron.log 2>&1
echo "$(date): Collector finished with exit code $?" >> cron.log
API Reference
The web interface exposes a REST API. FastAPI also serves interactive docs at /docs and OpenAPI at /openapi.json.
Configuration
| Method | Endpoint | Description |
|---|---|---|
| GET | /api/config |
Full config snapshot |
| GET | /api/subreddits |
List configured subreddits |
| POST | /api/subreddits |
Add a subreddit |
| DELETE | /api/subreddits/{name} |
Remove a subreddit |
| GET | /api/users |
List configured users |
| POST | /api/users |
Add a user |
| DELETE | /api/users/{name} |
Remove a user |
| GET | /api/settings |
Download + rate-limit settings |
| PUT | /api/settings |
Update settings |
Media
| Method | Endpoint | Description |
|---|---|---|
| GET | /api/media |
List downloaded media (paginated, filterable) |
| GET | /api/media/{id}/info |
Get media details |
| DELETE | /api/media/{id}?blacklist_author=&blacklist_subreddit= |
Delete media file (and optionally blacklist) |
| GET | /api/media/subreddits[?limit&offset] |
Subreddits with downloaded content |
| GET | /api/media/authors[?limit&offset] |
Authors with downloaded content |
| GET | /api/media/file/{filename:path} |
Serve raw file (Range-aware for video) |
| GET | /api/media/thumb/{filename:path} |
Serve / generate video thumbnail |
| GET | /api/media/blacklist-preview |
Files affected by current blacklist |
| POST | /api/media/cleanup-blacklist |
Delete blacklisted media |
| GET | /api/media/cleanup-preview?media_type= |
Files affected by type cleanup |
| POST | /api/media/cleanup-by-type?media_type= |
Delete by media type |
Blacklist
| Method | Endpoint | Description |
|---|---|---|
| GET | /api/blacklist |
Full blacklist |
| POST | /api/blacklist/{authors,subreddits,keywords,domains} |
Add entry |
| DELETE | /api/blacklist/{kind}/{name} |
Remove entry |
Favorites & Authors
| Method | Endpoint | Description |
|---|---|---|
| GET | /api/favorites |
Favorited posts (paginated) |
| POST | /api/favorites/{post_id} |
Add to favorites |
| DELETE | /api/favorites/{post_id} |
Remove from favorites |
| GET | /api/favorites/authors[?limit&offset] |
Distinct favorited authors |
| POST | /api/favorites/sync-users |
Add favorite authors as user targets |
| GET | /api/authors |
Authors with stats (paginated, sortable, favorites filter) |
| GET | /api/authors/{author}/media |
Media for one author |
Stats
| Method | Endpoint | Description |
|---|---|---|
| GET | /api/stats |
Disk + counts (30 s server-side cache) |
| GET | /api/stats/enhanced |
Trends, top authors, scores |
| GET | /api/stats/recent |
Recent downloads |
Collector & Scheduler
| Method | Endpoint | Description |
|---|---|---|
| POST | /api/collector/run |
Trigger collection run |
| GET | /api/collector/status |
Collector status |
| POST | /api/collect/individual |
Collect from a single subreddit/user |
| GET | /api/collect/targets |
List available targets |
| GET | /api/scheduler/status |
Scheduler state + next run |
| PUT | /api/scheduler/config |
Update schedule |
| GET | /api/scheduler/history |
Past scheduler runs |
| POST | /api/scheduler/run-now |
Execute schedule immediately |
Database Schema
The SQLite database (media.db) stores all metadata:
CREATE TABLE posts (
id TEXT PRIMARY KEY,
subreddit TEXT NOT NULL,
author TEXT,
title TEXT,
url TEXT NOT NULL,
media_url TEXT,
media_type TEXT,
score INTEGER DEFAULT 0,
created_utc REAL,
downloaded_at TIMESTAMP,
local_path TEXT,
file_hash TEXT,
permalink TEXT,
source_type TEXT,
flair TEXT
);
CREATE TABLE favorites (
post_id TEXT PRIMARY KEY,
favorited_at TIMESTAMP DEFAULT CURRENT_TIMESTAMP,
FOREIGN KEY (post_id) REFERENCES posts(id)
);
Troubleshooting
Videos saving as .html
Cause: yt-dlp not installed or not in PATH
Solution:
pip install yt-dlp
export PATH="$HOME/.local/bin:$PATH"
Rate limited (429 errors)
Cause: Too many requests to Reddit
Solution: Increase download_delay_seconds or decrease requests_per_minute in config
Incomplete galleries
Cause: Gallery metadata not available from Reddit API
Solution: Verify the post still exists on Reddit
Web interface not starting
Cause: Port already in use or missing dependencies
Solution:
# Check if port is in use
lsof -i :8000
# Reinstall dependencies (editable)
pip install -e ".[dev]"
Project Structure
reddit-media-collector/
├── src/
│ ├── main.py # Collector entry point
│ ├── config.py # Configuration dataclasses
│ ├── database.py # SQLite wrapper
│ ├── downloader.py # Downloader with retry + dedupe
│ ├── reddit_client.py # Reddit JSON-API client
│ ├── sidecar.py # Immich-compatible JSON sidecars
│ ├── extractors/ # URL extractors per host
│ │ ├── reddit.py
│ │ ├── imgur.py
│ │ └── gfycat.py
│ └── web/ # FastAPI app
│ ├── app.py
│ ├── auth.py # Optional HTTP Basic auth
│ ├── config_manager.py
│ ├── deps.py
│ ├── rate_limit.py # Per-IP throttle dependency
│ ├── routers/
│ │ ├── config.py
│ │ ├── favorites.py
│ │ ├── media.py
│ │ ├── scheduler.py
│ │ └── stats.py
│ ├── static/
│ │ └── js/api.js # Shared frontend helpers
│ └── templates/
│ └── index.html # SPA shell
├── tests/ # pytest (unit + API contract)
├── downloads/ # Downloaded media (gitignored)
├── config.yaml # Configuration
├── media.db # SQLite database (gitignored)
├── pyproject.toml # Project metadata + tooling config
├── .pre-commit-config.yaml # ruff + mypy hooks
├── .github/workflows/ci.yml # lint, types, test, docker
├── Dockerfile
├── docker-compose.yml
└── README.md
Contributing
Contributions are welcome! Please feel free to submit a Pull Request.
- Fork the repository
- Create your feature branch (
git checkout -b feature/amazing-feature) - Commit your changes (
git commit -m 'Add amazing feature') - Push to the branch (
git push origin feature/amazing-feature) - Open a Pull Request
License
This project is licensed under the MIT License - see the LICENSE file for details.
Disclaimer
This tool is for personal use only. Please respect Reddit's Terms of Service and the content creators' rights. Do not use this tool to redistribute copyrighted content.