project 01/02
PlateTracker.
Real-time license plate recognition for Croatian plates: YOLOv8 detection, object tracking and OCR that waits for agreement before it trusts a read.
- role
- Solo · design and implementation
- period
- 2025
- category
- Computer vision
- status
- Prototype
Private repository. Happy to walk through it on a call.
// overview
PlateTracker reads license plates from live video. A custom-trained YOLOv8 model finds plates frame by frame, every plate gets a persistent track, and an OCR stage reads the characters. A plate only counts once several reads agree. A second iteration, LicenceRIP, swapped the OCR engine for FastALPR configured for European and Croatian plates and compared the two.
- Python
- OpenCV
- YOLOv8
- PaddleOCR
- FastALPR
- NumPy
// problem
OCR on a single video frame is noisy. Plates are small, blurred and seen at an angle, and one bad frame turns ZG 1234-AB into ZG 1Z34-A8. Logging every raw read would bury the real plates under near-duplicates and errors, so the system has to decide when a reading is actually trustworthy, and it has to do it fast enough to keep up with the video.
// approach
- 01
Detect
A YOLOv8 model trained on plate data finds plates in every frame. Detections below a minimum pixel size are skipped, since they are too small to read reliably anyway.
- 02
Track
Each plate gets a track ID, so readings from different frames are grouped by vehicle instead of treated independently. History is capped at 120 frames per track to bound memory.
- 03
Read
OCR runs every second frame per track instead of every frame. Crops are trimmed by a small margin and upscaled 3× first, so distant plates still have enough pixels for the OCR to work with.
- 04
Confirm
Candidate reads are voted on across frames. A plate is confirmed after two consistent reads, or straight away at 0.8 confidence or more, and reads within one character of each other are merged into the same candidate.
- 05
Log
Confirmed plates are written once to a timestamped log. A debug view shows every crop the OCR sees, which made tuning the thresholds against real footage much faster.
// highlights
- Two OCR backends compared on the same pipeline: PaddleOCR (angle classification, GPU) and FastALPR tuned for EU/HR plates
- Multi-frame voting with a one-character similarity threshold to suppress OCR noise
- OCR throttled to every second frame per track to keep the loop real-time
- Every threshold in one config block: crop margin, upscale factor, confidence, history length
// what I learned
- Most of the accuracy came from what happens around the model (cropping, upscaling, voting), not from the model itself.
- Tracking turns a per-frame problem into a per-object one, and that is what makes reasoning over time possible.
- Making every threshold explicit config made tuning a matter of minutes instead of code changes.