Make result structs sendable for ASR - #46
Conversation
Speaker Diarization Benchmark ResultsSpeaker Diarization PerformanceEvaluating "who spoke when" detection accuracy
Diarization Pipeline Timing BreakdownTime spent in each stage of speaker diarization
Speaker Diarization Research ComparisonResearch baselines typically achieve 18-30% DER on standard datasets
🎯 Speaker Diarization Test • AMI Corpus ES2004a • 1049.4s meeting audio • 52.2s diarization time • Test runtime: 1m 21s • 07/29/2025, 07:22 PM EST |
VAD Benchmark ResultsPerformance Comparison
Dataset Details
✅ EXCELLENT: Average F1-Score above 70% |
ASR Benchmark Results
500 files per dataset • Test runtime: 42m36s • 07/29/2025, 08:04 PM EST RTFx = Real-Time Factor (higher is better) • Calculated as: Total audio duration ÷ Total processing time Expected RTFx Performance on Physical M1 Hardware:• M1 Mac: ~28x (clean), ~25x (other) Testing methodology follows HuggingFace Open ASR Leaderboard |
No description provided.