initial commit

This commit is contained in:
TylerCG 2026-08-20 09:38:30 -04:00
commit f8c8442a6e
4 changed files with 1553 additions and 0 deletions

651
README.md Normal file
View File

@ -0,0 +1,651 @@
# Plex Media Library Statistics Generator
A production-quality Python 3 command-line tool for scanning Plex-style media libraries and generating CSV statistics about storage size versus total video duration.
## Purpose
This tool analyzes media libraries (TV shows, anime, and movies) without requiring Plex itself or the Plex API. It operates directly on the filesystem and generates CSV reports showing:
- Storage size (in bytes and GB)
- Total video duration (in seconds and hours)
- Storage efficiency (GB per hour of video)
The key innovation is **incremental scanning with SQLite caching**: after the first full scan, subsequent scans are fast because unchanged files are never re-probed with FFmpeg.
## Features
- ✅ Supports TV shows, anime (treated as TV), and movies
- ✅ Automatic library type detection (override with `--type`)
- ✅ Persistent SQLite cache with change detection
- ✅ Fast incremental scans using file size + mtime signatures
- ✅ Comprehensive error handling (one bad file doesn't break the scan)
- ✅ Progress output during scanning and probing
- ✅ CSV output sorted alphabetically by series/movie
- ✅ Support for multiple libraries in a single database
- ✅ Per-video metadata stored for future feature expansion
## Requirements
### System
- Linux (though Windows/macOS paths would work with minor adjustments)
- Python 3.10+
- FFmpeg with `ffprobe` installed
### Python
Only Python standard library is used:
- `argparse` (CLI argument parsing)
- `pathlib` (filesystem operations)
- `sqlite3` (caching)
- `subprocess` (invoking ffprobe)
- `csv` (output generation)
## Installation
### Ubuntu / Debian
```bash
# Install FFmpeg
sudo apt update
sudo apt install ffmpeg
# Clone or download the tool
cd /path/to/media-scrapper
# Make the script executable (optional)
chmod +x plex_stats.py
# Test that ffprobe works
ffprobe -version
```
### CentOS / RHEL
```bash
sudo yum install ffmpeg
```
### macOS (using Homebrew)
```bash
brew install ffmpeg
```
## Quick Start
### Scan a TV library
```bash
python3 plex_stats.py /plex/tv --type tv --csv-prefix tv
```
Output:
- `tv_seasons.csv` — one row per season
- `tv_series.csv` — one row per complete series
- `plex_stats.db` — SQLite cache
### Scan an anime library
```bash
python3 plex_stats.py /plex/anime --type tv --csv-prefix anime
```
TV and anime use the same directory structure and CSV format.
### Scan a movies library
```bash
python3 plex_stats.py /plex/movies --type movies --csv-prefix movies
```
Output:
- `movies_movies.csv` — one row per movie folder
### Re-scan (fast)
Run the same command again. Files that haven't changed will be cached and won't be re-probed:
```bash
python3 plex_stats.py /plex/tv --type tv --csv-prefix tv
```
Expected output:
```
Scanning /plex/tv...
Files discovered: 4,812
Cached/unchanged: 4,790
New: 17
Changed: 5
Removed: 3
Probing 22 files...
[1/22] Breaking Bad...
[2/22] Game of Thrones...
...
Scan complete.
TV shows: 163
Seasons: 691
Video files: 4,809
Total size: 6.82 TB
Total runtime: 3,921.4 hours
Unchanged: 4,790
New: 14
Changed: 5
Removed: 3
Errors: 0
CSV files written:
tv_seasons.csv
tv_series.csv
Database:
plex_stats.db
```
## CLI Options
```
usage: plex_stats.py [-h] [--type {tv,movies}] [--db DB]
[--output-dir OUTPUT_DIR] [--csv-prefix CSV_PREFIX]
[--force] [--verbose]
library_root
positional arguments:
library_root Path to media library root (e.g. /plex/tv)
optional arguments:
-h, --help show help message
--type {tv,movies} Library type. If omitted, inferred from directory name.
--db DB Path to SQLite database (default: plex_stats.db)
--output-dir OUTPUT Directory for CSV files (default: current directory)
--csv-prefix PREFIX Prefix for CSV filenames (default: derived from root name)
--force Force ffprobe scan of all files (ignore cache)
--verbose Print verbose output
```
## Usage Examples
### Basic single-library scan
```bash
python3 plex_stats.py /plex/tv
```
Infers type as "tv", creates `plex_stats.db` in current directory, outputs `tv_seasons.csv` and `tv_series.csv`.
### All three libraries with explicit prefixes
```bash
python3 plex_stats.py /plex/tv --type tv --csv-prefix tv
python3 plex_stats.py /plex/anime --type tv --csv-prefix anime
python3 plex_stats.py /plex/movies --type movies --csv-prefix movies
```
All use the same `plex_stats.db`, creating six CSV files:
- `tv_seasons.csv`, `tv_series.csv`
- `anime_seasons.csv`, `anime_series.csv`
- `movies_movies.csv`
### Centralized database and output directory
```bash
python3 plex_stats.py /plex/tv \
--type tv \
--db /var/lib/plex-stats/plex_stats.db \
--output-dir /home/user/plex-reports \
--csv-prefix tv
```
### Force complete re-scan (e.g., after fixing bad ffprobe issues)
```bash
python3 plex_stats.py /plex/tv --force
```
Every file is re-probed. Useful if you suspect cache corruption or want to update all durations.
### Verbose output
```bash
python3 plex_stats.py /plex/tv --verbose
```
## Directory Structure
### TV / Anime
```
/plex/tv/
├── Breaking Bad/ <- Series name (first directory level)
│ ├── Season 01/ <- Season folder (second level)
│ │ ├── Breaking Bad - S01E01.mkv
│ │ ├── Breaking Bad - S01E02.mkv
│ │ └── ...
│ ├── Season 02/
│ │ └── ...
│ └── Specials/ <- Any season name is recognized
├── Game of Thrones/
│ ├── Season 01/
│ │ └── ...
│ └── Season 02/
│ └── ...
```
**Season folder names are flexible:**
- `Season 01`, `Season 1`
- `Season 00` (for pilots/specials)
- `Specials`, `Special`
- `S01`, `S1`
The actual folder name found on disk is preserved in the CSV output.
### Movies
```
/plex/movies/
├── Dune (2021)/ <- Movie name (first directory level)
│ └── Dune (2021).mkv
├── Interstellar (2014)/
│ ├── Interstellar (2014).mkv
│ └── Interstellar (2014) - Extras.mkv <- Multiple files per movie are aggregated
```
## Recognized Video Extensions
The tool recognizes these file extensions (case-insensitive):
- `.mkv`, `.mp4`, `.m4v`, `.avi`, `.mov`
- `.ts`, `.m2ts`, `.webm`
- `.mpg`, `.mpeg`, `.wmv`
Other files (subtitles, `.srt`, `.sub`, `.idx`, `.nfo`, images, metadata) are ignored.
## CSV Output Format
### TV Seasons (`*_seasons.csv`)
One row per season of a TV series:
```csv
Library,Series,Season,File_Count,Size_Bytes,Size_GB,Duration_Seconds,Duration_Hours,GB_Per_Hour
tv,Breaking Bad,Season 01,7,13786845020,12.84,20952,5.82,2.21
tv,Breaking Bad,Season 02,13,25887942819,24.11,38628,10.73,2.25
tv,Breaking Bad,Season 03,13,26943589405,25.10,38628,10.73,2.34
tv,Game of Thrones,Season 01,10,47129348901,43.91,41932,11.65,3.77
tv,Game of Thrones,Season 02,10,45219374812,42.13,41932,11.65,3.61
```
- **Library**: Always "tv" for TV/anime
- **Series**: Show name
- **Season**: Season folder name as found on disk
- **File_Count**: Number of video files in this season
- **Size_Bytes**: Total bytes (numeric, for calculation)
- **Size_GB**: Decimal gigabytes (1 GB = 1,000,000,000 bytes)
- **Duration_Seconds**: Total duration in seconds
- **Duration_Hours**: Total duration in hours
- **GB_Per_Hour**: Storage efficiency (size / duration)
### TV Series (`*_series.csv`)
One row per complete series, summing all seasons:
```csv
Library,Series,Season_Count,File_Count,Size_Bytes,Size_GB,Duration_Seconds,Duration_Hours,GB_Per_Hour
tv,Breaking Bad,5,62,128608176245,119.88,191500,53.19,2.25
tv,Game of Thrones,8,80,434328947123,404.56,334656,92.96,4.35
```
- **Season_Count**: Number of seasons in the series
- Other fields aggregate across all seasons
### Movies (`*_movies.csv`)
One row per movie:
```csv
Library,Movie,File_Count,Size_Bytes,Size_GB,Duration_Seconds,Duration_Hours,GB_Per_Hour
movies,Dune (2021),1,4829348901,4.50,10680,2.97,1.51
movies,Interstellar (2014),2,6234919234,5.81,14400,4.00,1.45
```
- **Movie**: Movie folder name
- **File_Count**: Number of video files under the movie folder (includes extras)
- Other fields aggregate all files
### CSV Import to Google Sheets
All numeric columns remain numeric (not formatted as strings), so they import cleanly into Google Sheets:
1. Open Google Sheets
2. Click "File" > "Import" > "Upload"
3. Select the CSV file
4. Choose "Replace spreadsheet"
5. Numeric columns are automatically detected
## SQLite Cache Database
### Schema
The tool automatically creates and maintains a SQLite database with this schema:
```sql
CREATE TABLE media_files (
id INTEGER PRIMARY KEY AUTOINCREMENT,
path TEXT UNIQUE NOT NULL,
library_root TEXT NOT NULL,
library_type TEXT NOT NULL,
show_movie TEXT NOT NULL,
season TEXT,
size_bytes INTEGER NOT NULL,
mtime_ns INTEGER NOT NULL,
duration_seconds REAL,
scan_timestamp REAL NOT NULL,
error_message TEXT
);
CREATE INDEX idx_library_root ON media_files(library_root);
CREATE INDEX idx_path ON media_files(path);
CREATE INDEX idx_show_movie ON media_files(show_movie);
CREATE INDEX idx_season ON media_files(season);
```
### Per-Video Storage
Every video file is stored individually, not just aggregated season/movie totals. This enables future features like:
- Per-episode statistics
- Codec and resolution analysis
- Average file size per series
- Duplicate detection
- Video quality metrics
## Change Detection & Incremental Scanning
### How It Works
1. **File discovery**: Walk the filesystem and list all video files.
2. **Stat check**: For each file, read `stat()` to get size and modification time.
3. **Cache lookup**: Check SQLite for this file's cached entry.
4. **Decision logic**:
- If cached entry exists **AND** size and mtime match → **Reuse cache** (no ffprobe)
- If cached entry exists **BUT** size or mtime differs → **Re-probe** (file changed)
- If no cached entry → **Probe** (new file)
- If cached entry is in SQLite **BUT** file no longer exists → **Delete from cache**
### Why Size + mtime?
- **Size**: Detects if video content changed
- **mtime (modification time)**: Detects if metadata or the file was touched
- **Together**: Extremely fast check without hashing or reading file content
- **Safe**: Works reliably for normal media workflows (copying files, minor edits)
### Limitations
If you manually edit file bytes without updating mtime (unusual), the cache will not detect it. This is acceptable for media libraries.
## Handling Unchanged, New, and Changed Files
### Unchanged
```
ffprobe not called → instant retrieval from cache
```
### New
```
File doesn't exist in cache → ffprobe called → database entry created
```
### Changed
```
File exists but size_bytes or mtime_ns differs → ffprobe called → database entry updated
```
### Removed
```
Cached file path no longer exists → removed from database
Note: Only scanned library root is checked.
Other libraries' files are not affected.
```
## Forced Rescans
Use `--force` to ignore the cache and re-probe every file:
```bash
python3 plex_stats.py /plex/tv --force
```
This is useful if:
- You suspect cache corruption
- FFmpeg was updated and you want fresh probes
- You want to update all durations to the latest ffprobe version
Forced rescans still respect the database schema and update all cached entries.
## Error Handling
### FFprobe Failures
If ffprobe fails on a specific file:
- A warning is printed to console
- The error is recorded in the database (`error_message` column)
- The scan continues with other files
- The file is eligible for retry on future scans
### What's Not an Error
- Permission denied reading a file → warning, continue
- Corrupted video file → ffprobe fails, recorded, continue
- Network timeouts on NFS → handled, continue
### Error Summary
At scan completion, total error count is displayed:
```
Errors: 5
```
To see details, query the database:
```bash
sqlite3 plex_stats.db "SELECT path, error_message FROM media_files WHERE error_message IS NOT NULL;"
```
## Database Management
### View cached files for a library
```bash
sqlite3 plex_stats.db "SELECT COUNT(*) FROM media_files WHERE library_root = '/plex/tv';"
```
### Export all files for a series
```bash
sqlite3 plex_stats.db "SELECT path, size_bytes, duration_seconds FROM media_files WHERE show_movie = 'Breaking Bad';"
```
### Clear all cache for one library
```bash
sqlite3 plex_stats.db "DELETE FROM media_files WHERE library_root = '/plex/tv';"
sqlite3 plex_stats.db "VACUUM;" # Reclaim disk space
```
### Clear entire database
```bash
rm plex_stats.db
```
The tool will recreate it on the next run.
## Performance Tips
1. **First scan is slow**: Thousands of ffprobe calls can take 1030 minutes for large libraries.
2. **Subsequent scans are fast**: Usually under a minute if only a few files changed.
3. **Use `--force` sparingly**: Only when you have a specific reason.
4. **Database is small**: Even for 10,000+ files, SQLite database is typically < 1 MB.
5. **CSV generation is instant**: All heavy lifting is in the scanning phase.
### Estimated Timeline
- 1,000 files: 310 minutes (first scan), 1030 seconds (subsequent)
- 5,000 files: 1550 minutes (first scan), 3060 seconds (subsequent)
- 10,000 files: 3090 minutes (first scan), 12 minutes (subsequent)
Varies based on filesystem speed, network latency (NFS/SMB), and average file size.
## Multiple Libraries in One Database
All three examples below use the same `plex_stats.db`:
```bash
python3 plex_stats.py /plex/tv --type tv --csv-prefix tv
python3 plex_stats.py /plex/anime --type tv --csv-prefix anime
python3 plex_stats.py /plex/movies --type movies --csv-prefix movies
```
**Database records are scoped by `library_root`**, so:
- Scanning `/plex/tv` only touches records where `library_root = '/plex/tv'`
- Deletion detection only removes files that belong to that library
- Each library can be scanned independently
This allows a single CSV output directory with files from multiple sources:
```
tv_seasons.csv (from /plex/tv)
tv_series.csv (from /plex/tv)
anime_seasons.csv (from /plex/anime)
anime_series.csv (from /plex/anime)
movies_movies.csv (from /plex/movies)
```
## Troubleshooting
### ffprobe not found
```
Error: ffprobe was not found.
Install FFmpeg, for example on Ubuntu/Debian:
sudo apt install ffmpeg
```
**Solution**: Install FFmpeg as shown in the error message.
### Permission denied scanning directory
```
Warning: Could not access /plex/tv: Permission denied
```
**Solution**: Run with appropriate permissions:
```bash
sudo python3 plex_stats.py /plex/tv
```
### Database locked
```
sqlite3.OperationalError: database is locked
```
**Solution**: Ensure only one instance of the tool is running. If multiple processes try to write simultaneously, wait for the first to finish.
### CSV files not created
Check that the output directory exists and is writable:
```bash
python3 plex_stats.py /plex/tv --output-dir /tmp
```
### Strange results in CSV
1. Verify the directory structure matches TV or movie format
2. Check for videos in unexpected locations
3. Re-scan with `--force` to regenerate all probes
## Extending the Tool
The database stores per-video metadata, so future enhancements are straightforward:
### Adding codec analysis
```python
# Expand ffprobe query to include codec
ffprobe ... -show_entries stream=codec_name
# Add columns: video_codec, audio_codec
# Aggregate in CSV generation
```
### Adding resolution tracking
```python
# Query video stream height
ffprobe ... -show_entries stream=height
# Add columns: resolution, resolution_count
```
### Per-episode breakdown
```python
# Parse episode numbers from filenames
# Create *_episodes.csv with per-episode stats
```
All of this is possible without redesigning the cache because per-video data is preserved.
## License
MIT License. Use freely.
## Support
For issues or feature requests, consult the source code and modify as needed. The tool is designed to be maintainable and extensible.
---
**Example complete workflow:**
```bash
#!/bin/bash
# Install FFmpeg (first time only)
sudo apt install ffmpeg
# Create output directory
mkdir -p ~/plex-reports
# Scan TV library
python3 plex_stats.py /plex/tv \
--type tv \
--db ~/plex-stats.db \
--output-dir ~/plex-reports \
--csv-prefix tv
# Scan anime library
python3 plex_stats.py /plex/anime \
--type tv \
--db ~/plex-stats.db \
--output-dir ~/plex-reports \
--csv-prefix anime
# Scan movies library
python3 plex_stats.py /plex/movies \
--type movies \
--db ~/plex-stats.db \
--output-dir ~/plex-reports \
--csv-prefix movies
# View results
ls -lh ~/plex-reports/*.csv
# Import into Google Sheets:
# - Open Google Sheets
# - File > Import
# - Upload ~/plex-reports/tv_seasons.csv
```

BIN
plex_stats.db Normal file

Binary file not shown.

863
plex_stats.py Normal file
View File

@ -0,0 +1,863 @@
#!/usr/bin/env python3
"""
Plex Media Library Statistics Generator
Scans a Plex-style media library and generates CSV statistics about movies, TV shows, and anime.
Uses SQLite for persistent caching to avoid rescanning unchanged files.
Author: Media Scraper Tool
License: MIT
"""
import argparse
import csv
import os
import sqlite3
import subprocess
import sys
import time
from dataclasses import dataclass
from datetime import datetime
from pathlib import Path
from typing import Dict, List, Optional, Set, Tuple
# Configuration constants
VIDEO_EXTENSIONS = {'.mkv', '.mp4', '.m4v', '.avi', '.mov', '.ts', '.m2ts', '.webm', '.mpg', '.mpeg', '.wmv'}
BYTES_PER_GB = 1_000_000_000
SECONDS_PER_HOUR = 3600
CSV_DECIMAL_PLACES = 2
@dataclass
class MediaFile:
"""Represents a single video file with its metadata."""
path: str
library_root: str
library_type: str # 'tv' or 'movies'
show_movie: str
season: Optional[str] # None for movies
size_bytes: int
mtime_ns: int
duration_seconds: Optional[float]
scan_timestamp: float
error_message: Optional[str] = None
class PlexStatsDB:
"""Manages SQLite database for caching media file metadata."""
def __init__(self, db_path: Path):
"""Initialize database connection and create schema if needed."""
self.db_path = db_path
self.conn: Optional[sqlite3.Connection] = None
self._init_db()
def _init_db(self) -> None:
"""Initialize database connection and create tables if needed."""
self.conn = sqlite3.connect(self.db_path)
self.conn.row_factory = sqlite3.Row
cursor = self.conn.cursor()
# Create media_files table
cursor.execute('''
CREATE TABLE IF NOT EXISTS media_files (
id INTEGER PRIMARY KEY AUTOINCREMENT,
path TEXT UNIQUE NOT NULL,
library_root TEXT NOT NULL,
library_type TEXT NOT NULL,
show_movie TEXT NOT NULL,
season TEXT,
size_bytes INTEGER NOT NULL,
mtime_ns INTEGER NOT NULL,
duration_seconds REAL,
scan_timestamp REAL NOT NULL,
error_message TEXT
)
''')
# Create indexes for common queries
cursor.execute('CREATE INDEX IF NOT EXISTS idx_library_root ON media_files(library_root)')
cursor.execute('CREATE INDEX IF NOT EXISTS idx_path ON media_files(path)')
cursor.execute('CREATE INDEX IF NOT EXISTS idx_show_movie ON media_files(show_movie)')
cursor.execute('CREATE INDEX IF NOT EXISTS idx_season ON media_files(season)')
self.conn.commit()
def close(self) -> None:
"""Close database connection."""
if self.conn:
self.conn.close()
def __enter__(self):
"""Context manager entry."""
return self
def __exit__(self, exc_type, exc_val, exc_tb):
"""Context manager exit."""
self.close()
def get_cached_file(self, path: str) -> Optional[MediaFile]:
"""Retrieve a cached media file by path."""
cursor = self.conn.cursor()
cursor.execute('SELECT * FROM media_files WHERE path = ?', (path,))
row = cursor.fetchone()
if not row:
return None
return MediaFile(
path=row['path'],
library_root=row['library_root'],
library_type=row['library_type'],
show_movie=row['show_movie'],
season=row['season'],
size_bytes=row['size_bytes'],
mtime_ns=row['mtime_ns'],
duration_seconds=row['duration_seconds'],
scan_timestamp=row['scan_timestamp'],
error_message=row['error_message']
)
def insert_or_update_file(self, media_file: MediaFile) -> None:
"""Insert or update a media file record."""
cursor = self.conn.cursor()
cursor.execute('''
INSERT OR REPLACE INTO media_files
(path, library_root, library_type, show_movie, season, size_bytes,
mtime_ns, duration_seconds, scan_timestamp, error_message)
VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?)
''', (
media_file.path,
media_file.library_root,
media_file.library_type,
media_file.show_movie,
media_file.season,
media_file.size_bytes,
media_file.mtime_ns,
media_file.duration_seconds,
media_file.scan_timestamp,
media_file.error_message
))
self.conn.commit()
def get_files_by_library_root(self, library_root: str) -> List[MediaFile]:
"""Retrieve all files cached for a specific library root."""
cursor = self.conn.cursor()
cursor.execute('SELECT * FROM media_files WHERE library_root = ?', (library_root,))
rows = cursor.fetchall()
return [
MediaFile(
path=row['path'],
library_root=row['library_root'],
library_type=row['library_type'],
show_movie=row['show_movie'],
season=row['season'],
size_bytes=row['size_bytes'],
mtime_ns=row['mtime_ns'],
duration_seconds=row['duration_seconds'],
scan_timestamp=row['scan_timestamp'],
error_message=row['error_message']
)
for row in rows
]
def delete_file(self, path: str) -> None:
"""Delete a file record from the database."""
cursor = self.conn.cursor()
cursor.execute('DELETE FROM media_files WHERE path = ?', (path,))
self.conn.commit()
def check_ffprobe_available() -> bool:
"""Check if ffprobe is available in PATH."""
try:
subprocess.run(
['ffprobe', '-version'],
capture_output=True,
timeout=5,
check=True
)
return True
except (subprocess.CalledProcessError, FileNotFoundError, subprocess.TimeoutExpired):
return False
def get_duration_from_ffprobe(file_path: str) -> Optional[float]:
"""
Extract duration in seconds from a video file using ffprobe.
Returns:
Duration in seconds as float, or None if extraction fails.
"""
try:
result = subprocess.run(
[
'ffprobe',
'-v', 'error',
'-show_entries', 'format=duration',
'-of', 'default=noprint_wrappers=1:nokey=1',
file_path
],
capture_output=True,
timeout=30,
text=True,
check=False
)
if result.returncode != 0:
return None
duration_str = result.stdout.strip()
if not duration_str:
return None
return float(duration_str)
except (subprocess.TimeoutExpired, ValueError, OSError) as e:
return None
def infer_library_type(library_root: str) -> str:
"""
Infer library type from directory name.
Returns 'tv' or 'movies' based on directory name.
Defaults to 'tv' if uncertain.
"""
root_name = Path(library_root).name.lower()
if 'movie' in root_name or 'film' in root_name:
return 'movies'
return 'tv'
def extract_show_movie_name(file_path: str, library_root: str) -> Tuple[str, Optional[str]]:
"""
Extract show/movie name and season (if applicable) from file path.
Args:
file_path: Full path to the video file
library_root: Root of the library
Returns:
Tuple of (show/movie_name, season_name)
season_name is None for movies.
"""
relative_path = Path(file_path).relative_to(library_root)
parts = relative_path.parts
if len(parts) < 2:
# Unusual structure, use filename
return file_path.split(os.sep)[-2], None
show_movie = parts[0]
season = parts[1] if len(parts) > 2 else None
return show_movie, season
def get_video_files_in_library(library_root: Path, library_type: str) -> List[str]:
"""
Recursively find all video files in a library root.
Args:
library_root: Path to library root directory
library_type: 'tv' or 'movies'
Returns:
List of absolute paths to video files.
"""
video_files: List[str] = []
try:
for path in library_root.rglob('*'):
if path.is_file() and path.suffix.lower() in VIDEO_EXTENSIONS:
video_files.append(str(path.resolve()))
except (PermissionError, OSError) as e:
print(f"Warning: Could not access {library_root}: {e}")
return video_files
def scan_library(
library_root: Path,
library_type: str,
db: PlexStatsDB,
force_rescan: bool = False,
verbose: bool = False
) -> Tuple[List[MediaFile], Dict[str, int]]:
"""
Scan a media library and update cache.
Args:
library_root: Path to library root
library_type: 'tv' or 'movies'
db: Database instance
force_rescan: If True, ffprobe all files regardless of cache
verbose: Print verbose output
Returns:
Tuple of (list of MediaFile objects, stats dict)
"""
library_root_str = str(library_root.resolve())
stats = {
'discovered': 0,
'cached_unchanged': 0,
'new': 0,
'changed': 0,
'removed': 0,
'errors': 0
}
print(f"\nScanning {library_root_str}...")
# Get all video files currently in library
current_files = get_video_files_in_library(library_root, library_type)
current_files_set: Set[str] = set(current_files)
stats['discovered'] = len(current_files)
print(f"Files discovered: {stats['discovered']:,}")
# Get previously cached files for this library
cached_files = db.get_files_by_library_root(library_root_str)
cached_by_path: Dict[str, MediaFile] = {f.path: f for f in cached_files}
# Find removed files
for cached_path in cached_by_path:
if cached_path not in current_files_set:
db.delete_file(cached_path)
stats['removed'] += 1
# Process current files
media_files: List[MediaFile] = []
files_to_probe: List[str] = []
for file_path in current_files:
try:
stat_info = os.stat(file_path)
size_bytes = stat_info.st_size
mtime_ns = stat_info.st_mtime_ns
show_movie, season = extract_show_movie_name(file_path, library_root_str)
# Check cache
cached = cached_by_path.get(file_path)
if (
not force_rescan
and cached is not None
and cached.size_bytes == size_bytes
and cached.mtime_ns == mtime_ns
):
# File unchanged, reuse cache
media_files.append(cached)
stats['cached_unchanged'] += 1
else:
# File is new or changed
media_file = MediaFile(
path=file_path,
library_root=library_root_str,
library_type=library_type,
show_movie=show_movie,
season=season,
size_bytes=size_bytes,
mtime_ns=mtime_ns,
duration_seconds=None,
scan_timestamp=time.time()
)
if cached is not None:
stats['changed'] += 1
else:
stats['new'] += 1
files_to_probe.append(file_path)
media_files.append(media_file)
except (OSError, PermissionError) as e:
print(f"Warning: Could not access {file_path}: {e}")
stats['errors'] += 1
print(f"Cached/unchanged: {stats['cached_unchanged']:,}")
print(f"New: {stats['new']}")
print(f"Changed: {stats['changed']}")
print(f"Removed: {stats['removed']}")
# Probe new/changed files
if files_to_probe:
print(f"\nProbing {len(files_to_probe)} files...")
for idx, file_path in enumerate(files_to_probe, 1):
# Find the media file object for this path
media_file = next((mf for mf in media_files if mf.path == file_path), None)
if media_file is None:
continue
# Print progress
show_name = media_file.show_movie[:50]
print(f" [{idx}/{len(files_to_probe)}] {show_name}...", flush=True)
# Get duration
duration = get_duration_from_ffprobe(file_path)
if duration is not None:
media_file.duration_seconds = duration
else:
media_file.error_message = "ffprobe failed or returned no duration"
stats['errors'] += 1
# Update database
db.insert_or_update_file(media_file)
# Insert cached files that weren't probed
for media_file in media_files:
if media_file.path not in files_to_probe:
db.insert_or_update_file(media_file)
return media_files, stats
def aggregate_tv_seasons(media_files: List[MediaFile]) -> Dict[Tuple[str, str], Dict]:
"""
Aggregate media files by series and season for TV output.
Args:
media_files: List of MediaFile objects
Returns:
Dict mapping (series, season) tuples to aggregated stats
"""
seasons: Dict[Tuple[str, str], Dict] = {}
for media_file in media_files:
if media_file.season is None:
continue
key = (media_file.show_movie, media_file.season)
if key not in seasons:
seasons[key] = {
'series': media_file.show_movie,
'season': media_file.season,
'file_count': 0,
'size_bytes': 0,
'duration_seconds': 0.0,
'error_count': 0
}
seasons[key]['file_count'] += 1
seasons[key]['size_bytes'] += media_file.size_bytes
if media_file.duration_seconds is not None:
seasons[key]['duration_seconds'] += media_file.duration_seconds
if media_file.error_message:
seasons[key]['error_count'] += 1
return seasons
def aggregate_tv_series(season_aggregates: Dict[Tuple[str, str], Dict]) -> Dict[str, Dict]:
"""
Aggregate season data by series for series-level output.
Args:
season_aggregates: Dict from aggregate_tv_seasons
Returns:
Dict mapping series names to aggregated stats
"""
series: Dict[str, Dict] = {}
for (series_name, season_name), season_data in season_aggregates.items():
if series_name not in series:
series[series_name] = {
'series': series_name,
'season_count': 0,
'file_count': 0,
'size_bytes': 0,
'duration_seconds': 0.0
}
series[series_name]['season_count'] += 1
series[series_name]['file_count'] += season_data['file_count']
series[series_name]['size_bytes'] += season_data['size_bytes']
series[series_name]['duration_seconds'] += season_data['duration_seconds']
return series
def aggregate_movies(media_files: List[MediaFile]) -> Dict[str, Dict]:
"""
Aggregate media files by movie name.
Args:
media_files: List of MediaFile objects (should all be movies)
Returns:
Dict mapping movie names to aggregated stats
"""
movies: Dict[str, Dict] = {}
for media_file in media_files:
movie_name = media_file.show_movie
if movie_name not in movies:
movies[movie_name] = {
'movie': movie_name,
'file_count': 0,
'size_bytes': 0,
'duration_seconds': 0.0,
'error_count': 0
}
movies[movie_name]['file_count'] += 1
movies[movie_name]['size_bytes'] += media_file.size_bytes
if media_file.duration_seconds is not None:
movies[movie_name]['duration_seconds'] += media_file.duration_seconds
if media_file.error_message:
movies[movie_name]['error_count'] += 1
return movies
def format_number(value: float, decimal_places: int = CSV_DECIMAL_PLACES) -> float:
"""Format a number to a specific number of decimal places."""
return round(value, decimal_places)
def write_tv_season_csv(
season_aggregates: Dict[Tuple[str, str], Dict],
output_path: Path,
csv_prefix: str
) -> None:
"""Write season-level TV statistics to CSV."""
csv_file = output_path / f"{csv_prefix}_seasons.csv"
# Sort by series, then season
sorted_seasons = sorted(
season_aggregates.items(),
key=lambda x: (x[0][0], x[0][1]) # (series, season)
)
with open(csv_file, 'w', newline='') as f:
writer = csv.DictWriter(
f,
fieldnames=[
'Library', 'Series', 'Season', 'File_Count',
'Size_Bytes', 'Size_GB', 'Duration_Seconds', 'Duration_Hours', 'GB_Per_Hour'
]
)
writer.writeheader()
for (series, season), data in sorted_seasons:
size_gb = format_number(data['size_bytes'] / BYTES_PER_GB)
duration_hours = format_number(data['duration_seconds'] / SECONDS_PER_HOUR)
gb_per_hour = (
format_number(size_gb / duration_hours)
if duration_hours > 0 else 0.0
)
writer.writerow({
'Library': 'tv',
'Series': series,
'Season': season,
'File_Count': data['file_count'],
'Size_Bytes': data['size_bytes'],
'Size_GB': size_gb,
'Duration_Seconds': data['duration_seconds'],
'Duration_Hours': duration_hours,
'GB_Per_Hour': gb_per_hour
})
print(f" {csv_file}")
def write_tv_series_csv(
series_aggregates: Dict[str, Dict],
output_path: Path,
csv_prefix: str
) -> None:
"""Write series-level TV statistics to CSV."""
csv_file = output_path / f"{csv_prefix}_series.csv"
# Sort by series name
sorted_series = sorted(series_aggregates.items(), key=lambda x: x[0])
with open(csv_file, 'w', newline='') as f:
writer = csv.DictWriter(
f,
fieldnames=[
'Library', 'Series', 'Season_Count', 'File_Count',
'Size_Bytes', 'Size_GB', 'Duration_Seconds', 'Duration_Hours', 'GB_Per_Hour'
]
)
writer.writeheader()
for series_name, data in sorted_series:
size_gb = format_number(data['size_bytes'] / BYTES_PER_GB)
duration_hours = format_number(data['duration_seconds'] / SECONDS_PER_HOUR)
gb_per_hour = (
format_number(size_gb / duration_hours)
if duration_hours > 0 else 0.0
)
writer.writerow({
'Library': 'tv',
'Series': series_name,
'Season_Count': data['season_count'],
'File_Count': data['file_count'],
'Size_Bytes': data['size_bytes'],
'Size_GB': size_gb,
'Duration_Seconds': data['duration_seconds'],
'Duration_Hours': duration_hours,
'GB_Per_Hour': gb_per_hour
})
print(f" {csv_file}")
def write_movies_csv(
movies_aggregates: Dict[str, Dict],
output_path: Path,
csv_prefix: str
) -> None:
"""Write movie statistics to CSV."""
csv_file = output_path / f"{csv_prefix}_movies.csv"
# Sort by movie name
sorted_movies = sorted(movies_aggregates.items(), key=lambda x: x[0])
with open(csv_file, 'w', newline='') as f:
writer = csv.DictWriter(
f,
fieldnames=[
'Library', 'Movie', 'File_Count',
'Size_Bytes', 'Size_GB', 'Duration_Seconds', 'Duration_Hours', 'GB_Per_Hour'
]
)
writer.writeheader()
for movie_name, data in sorted_movies:
size_gb = format_number(data['size_bytes'] / BYTES_PER_GB)
duration_hours = format_number(data['duration_seconds'] / SECONDS_PER_HOUR)
gb_per_hour = (
format_number(size_gb / duration_hours)
if duration_hours > 0 else 0.0
)
writer.writerow({
'Library': 'movies',
'Movie': movie_name,
'File_Count': data['file_count'],
'Size_Bytes': data['size_bytes'],
'Size_GB': size_gb,
'Duration_Seconds': data['duration_seconds'],
'Duration_Hours': duration_hours,
'GB_Per_Hour': gb_per_hour
})
print(f" {csv_file}")
def print_scan_summary(
library_type: str,
media_files: List[MediaFile],
stats: Dict[str, int],
season_agg: Optional[Dict[Tuple[str, str], Dict]] = None,
series_agg: Optional[Dict[str, Dict]] = None,
movie_agg: Optional[Dict[str, Dict]] = None,
db_path: Path = None,
csv_files: List[str] = None
) -> None:
"""Print a human-readable summary of the scan."""
if csv_files is None:
csv_files = []
# Calculate totals
total_size_gb = sum(mf.size_bytes for mf in media_files) / BYTES_PER_GB
total_duration_seconds = sum(
(mf.duration_seconds or 0) for mf in media_files if mf.duration_seconds
)
total_duration_hours = total_duration_seconds / SECONDS_PER_HOUR
print("\n" + "=" * 60)
print("Scan Complete")
print("=" * 60)
if library_type == 'tv':
unique_series = len(series_agg) if series_agg else 0
unique_seasons = len(season_agg) if season_agg else 0
print(f"TV series: {unique_series}")
print(f"Seasons: {unique_seasons}")
else:
unique_movies = len(movie_agg) if movie_agg else 0
print(f"Movies: {unique_movies}")
print(f"Video files: {len(media_files):,}")
print(f"Total size: {format_number(total_size_gb, 2)} GB")
print(f"Total runtime: {format_number(total_duration_hours, 1)} hours")
print()
print(f"Unchanged: {stats['cached_unchanged']:,}")
print(f"New: {stats['new']}")
print(f"Changed: {stats['changed']}")
print(f"Removed: {stats['removed']}")
print(f"Errors: {stats['errors']}")
if csv_files:
print()
print("CSV files written:")
for csv_file in csv_files:
print(f" {csv_file}")
if db_path:
print()
print("Database:")
print(f" {db_path}")
print("=" * 60)
def main():
"""Main entry point."""
parser = argparse.ArgumentParser(
description='Generate statistics for Plex-style media libraries.'
)
parser.add_argument(
'library_root',
help='Path to the media library root (e.g., /plex/tv or /plex/movies)'
)
parser.add_argument(
'--type',
choices=['tv', 'movies'],
help='Library type (tv or movies). If omitted, will be inferred from directory name.'
)
parser.add_argument(
'--db',
type=Path,
default=Path('plex_stats.db'),
help='Path to SQLite database (default: plex_stats.db)'
)
parser.add_argument(
'--output-dir',
type=Path,
default=Path('.'),
help='Directory for output CSV files (default: current directory)'
)
parser.add_argument(
'--csv-prefix',
help='Prefix for CSV filenames (e.g., "tv" -> tv_seasons.csv). '
'If omitted, derived from library directory name.'
)
parser.add_argument(
'--force',
action='store_true',
help='Force ffprobe scan of all files, ignoring cache'
)
parser.add_argument(
'--verbose',
action='store_true',
help='Print verbose output'
)
args = parser.parse_args()
# Validate inputs
library_root = Path(args.library_root).resolve()
if not library_root.exists():
print(f"Error: Library root does not exist: {library_root}")
sys.exit(1)
if not library_root.is_dir():
print(f"Error: Library root is not a directory: {library_root}")
sys.exit(1)
# Check ffprobe availability
if not check_ffprobe_available():
print("Error: ffprobe was not found.")
print()
print("Install FFmpeg, for example on Ubuntu/Debian:")
print()
print(" sudo apt install ffmpeg")
print()
sys.exit(1)
# Determine library type
library_type = args.type or infer_library_type(str(library_root))
# Determine CSV prefix
csv_prefix = args.csv_prefix or library_root.name
# Create output directory if needed
output_dir = args.output_dir
output_dir.mkdir(parents=True, exist_ok=True)
# Open database
with PlexStatsDB(args.db) as db:
# Scan library
media_files, stats = scan_library(
library_root,
library_type,
db,
force_rescan=args.force,
verbose=args.verbose
)
# Filter out files with errors for aggregation
valid_files = [mf for mf in media_files if mf.duration_seconds is not None]
csv_files_written: List[str] = []
# Generate CSV output
print("\nGenerating CSV files...")
if library_type == 'tv':
season_agg = aggregate_tv_seasons(valid_files)
series_agg = aggregate_tv_series(season_agg)
if season_agg:
write_tv_season_csv(season_agg, output_dir, csv_prefix)
csv_files_written.append(str(output_dir / f"{csv_prefix}_seasons.csv"))
if series_agg:
write_tv_series_csv(series_agg, output_dir, csv_prefix)
csv_files_written.append(str(output_dir / f"{csv_prefix}_series.csv"))
print_scan_summary(
library_type,
valid_files,
stats,
season_agg=season_agg,
series_agg=series_agg,
db_path=args.db,
csv_files=csv_files_written
)
else: # movies
movie_agg = aggregate_movies(valid_files)
if movie_agg:
write_movies_csv(movie_agg, output_dir, csv_prefix)
csv_files_written.append(str(output_dir / f"{csv_prefix}_movies.csv"))
print_scan_summary(
library_type,
valid_files,
stats,
movie_agg=movie_agg,
db_path=args.db,
csv_files=csv_files_written
)
if __name__ == '__main__':
main()

39
requirements.txt Normal file
View File

@ -0,0 +1,39 @@
# plex_stats.py has ZERO external Python dependencies.
#
# This tool uses only Python standard library:
# - argparse
# - csv
# - pathlib
# - sqlite3
# - subprocess
# - typing
#
# SYSTEM REQUIREMENTS (not Python packages):
# - Python 3.10 or later
# - FFmpeg with ffprobe (install via system package manager)
# Ubuntu/Debian: sudo apt install ffmpeg
# CentOS/RHEL: sudo yum install ffmpeg
# macOS: brew install ffmpeg
#
# To install FFmpeg on various systems:
#
# Ubuntu/Debian:
# sudo apt update && sudo apt install ffmpeg
#
# CentOS/RHEL:
# sudo yum install ffmpeg
#
# Fedora:
# sudo dnf install ffmpeg
#
# macOS (Homebrew):
# brew install ffmpeg
#
# Alpine Linux:
# apk add ffmpeg
#
# Verify installation:
# ffprobe -version
#
# This file is provided for reference only.
# No Python packages need to be installed via pip.