← All projects
Audio ML · Real-time

Gunfire Detection from Video

Streams audio from any video/RTSP feed and flags gunshots in under 100ms per window — no buffering delay.

FFmpeg Librosa TensorFlow scikit-learn

How it works

A prior version of this system buffered audio and ran a heavy image-style model per clip — accurate, but too slow for a live feed. This version streams small audio windows continuously through a lightweight classifier.

Data Sourcing

Combined gunshot recordings with broad real-world sound data for realistic positive and negative coverage.

Kaggle

Feature Extraction

Each short audio window is reduced to a compact spectral fingerprint — not a full spectrogram image.

Librosa · MFCC

Model Training

A lightweight classifier is trained and validated on a held-out split to check it generalises, not just memorises.

TensorFlow · scikit-learn

Real-Time Streaming

Audio is decoded continuously from video, RTSP, or mic input in small overlapping windows — nothing is buffered.

FFmpeg

Live Alerting

Each window is scored in roughly 70–90ms and alerts surface instantly with a confidence score.

Python
~80ms
Inference latency per window
0.5s
Detection lag from real gunshot onset
3
Source types: video file, RTSP, microphone

More projects