Harsh Wardhan
Identity: Dataset EngineeringResearch ToolDataset Pipeline

Reson Collector

Dataset Recording & Audio Labeling Pipeline for Acoustic ML

Engineering Challenge

“How do you build an acoustic dataset collection pipeline from scratch?”

AI•Completed•Creator•2026
RecordingAutomatedTimed Batches
ExportPCM WAVE16-bit 44.1kHz
JSON SchemaMetadataLabel Alignment
SanitizationZero ShiftClipping Detection
Reson Collector
01 Hypothesis

Research Problem & Motivation

Machine learning models for signal processing depend heavily on clean, consistently labeled datasets. Reson Collector automates the recording workflow, enforcing consistent distance, duration, and sampling parameters across hundreds of audio gesture trials.

Why I Built It

Manually recording, naming, and organizing audio clips for machine learning model training was slow and error-prone. I needed a tool that ensured every training sample had identical recording conditions and metadata.

02 Signal Processing

Signal & Data Pipeline

Pipeline Flow Architecture

Audio Prompt -> SoundDevice Recording Buffer -> Clipping/Silence Threshold Check -> WAV Encoder -> JSON Manifest Generator.

04 Technical Rigor

Engineering Decisions

Decision #01 • Direct PortAudio Buffering vs Disk Streaming
Problem:

Writing audio buffers to disk during active recording caused minor frame drops on lower-spec hardware.

Decision:

Buffered raw audio frames in memory NumPy arrays and flushed to disk asynchronously after trial completion.

Tradeoff:

Higher temporary RAM usage during long recording sessions.

Outcome:

Eliminated audio frame drops across all dataset recording sessions.

05 Obstacles

Challenges & Solutions

Detecting Clipped or Corrupted Trials

Issue: Users occasionally performed gestures outside the microphone range, producing useless silent audio samples.

Decision/Solution: Added instant post-trial amplitude threshold checks to automatically reject silent or clipped recordings before saving.

Future Roadmap

Future Research Directions

Automated spectrogram preview generation during capture
Cloud storage dataset sync
Multi-microphone channel recording support
Retrospective Summary

Key Takeaways

Takeaway #01

High-quality dataset collection pipelines save vast amounts of debugging time during ML model training.

Takeaway #02

Enforcing metadata structure at capture time prevents dataset inconsistency downstream.

Continue Exploring

Related Engineering Projects