Skip to content

feat: add dual-track recording to capture each side of a call separately - #72

Closed
R2bEEaton wants to merge 1 commit into
kitsumed:mainfrom
R2bEEaton:feature/dual-track-recording
Closed

R2bEEaton wants to merge 1 commit into
kitsumed:mainfrom
R2bEEaton:feature/dual-track-recording

Conversation

@R2bEEaton

Copy link
Copy Markdown

Summary

Adds a "Record each side separately" toggle (Settings -> Audio configuration) that captures the uplink (your mic) and downlink (the other party) as two independent, time-synced files instead of one mixed recording, e.g. call_..._uplink.opus + call_..._downlink.opus.

  • ShellService now runs a second, independent ShellAudioPipeline instance behind two new AIDL methods (startSecondaryRecording/stopSecondaryRecording), so a second scrcpy-server process (voice-call-downlink) runs concurrently alongside the existing one (voice-call-uplink), each with its own socket/pipe.
  • AudioRecordingEngine was reworked around a private TrackSession bundle so it can drive one (normal) or two (dual-track) capture pipelines in parallel, with symmetric start/rollback/teardown for both.
  • ScrcpyAudioMuxer accepts an optional shared wall-clock origin so both files' PTS=0 line up to the exact same instant, keeping the two tracks in sync for downstream processing.
  • RecordingFileNameFormatter gained an optional filename suffix (_uplink/_downlink) to distinguish the two outputs.
  • When the toggle is on, the Audio source dropdown is hidden (the uplink/downlink sources are forced) instead of shown disabled.

Why this matters

Today, recording both sides of a call only produces one mixed file (VOICE_CALL), or a single-sided file (VOICE_CALL_UPLINK/VOICE_CALL_DOWNLINK) if you pick one manually — there's no way to get both sides at once without them being mixed together. For transcription workflows, a mixed track makes speaker attribution unreliable (diarization heuristics), whereas two separate, perfectly-synced mono files make "who said what" deterministic: it's just "which file". This mirrors how call-center transcription products (e.g. Twilio/Deepgram dual-channel) typically separate speakers.

Two files (rather than one file with two audio tracks) was chosen deliberately: Android's MediaMuxer only supports a single audio track per OGG container, so a true single-file two-track approach would only work for AAC/MP4, not Opus. Two mono files work identically for both codecs, need no new audio pipeline stage, and can still be losslessly combined into one stereo file afterward with a single ffmpeg command (documented in docs/configuration.md) if a specific transcription API requires that shape instead.

Testing

  • Verified end-to-end on a real device (Pixel 10 Pro, Android 16) via Shizuku: placed and received real calls with the toggle off (unchanged single-file behavior, confirmed byte-identical code path) and on (confirmed two files appear in the recording folder with _uplink/_downlink suffixes, both playable, both non-empty, both starting from the same wall-clock instant).
  • Confirmed pause/resume and mid-call cancel/error paths correctly tear down both tracks with no orphaned scrcpy-server process and no partial/locked output file.
  • Confirmed toggling the setting on/off correctly hides/reveals the Audio source dropdown.
  • Full ./gradlew assembleDebug build passes with no new warnings from the changed files.

This is additive and off by default; existing single-track recording behavior is unchanged when the toggle is off.

Adds a "Record each side separately" toggle that captures the uplink
(your mic) and downlink (the other party) as two independent, time-synced
files instead of one mixed recording, so each speaker can be clearly
attributed for transcription.

Assisted-by: Claude
@R2bEEaton
R2bEEaton force-pushed the feature/dual-track-recording branch from 64423c7 to 5a67e77 Compare July 22, 2026 18:54
@R2bEEaton

Copy link
Copy Markdown
Author

Apologies — I should have opened an issue first to discuss this before submitting a PR, per the contributing guidelines (large features should be discussed before the work is done). Closing this for now; I'll open an issue instead so we can discuss the approach first.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant