Key Features
Phonexia Voice Inspector offers several features to effectively support voice forensic experts:
- Standalone application with a comprehensive, easy-to-use Graphical User Interface (GUI)
- Automatic speaker comparison of one or more questioned recordings (unknown-speaker recordings or voiceprints) against one or more suspected reference speakers, supporting 1:1, 1:N, and N:M comparison modes.
- Integrated Deepfake Detection to identify synthetic audio alongside speaker comparison results.
- Integrated speech technologies: Speaker Identification, Speaker Diarization, Phoneme Recognition, Voice Activity Detection, and Speech Quality Estimation.
- Ability to search for repetitive sound patterns across all recordings using automatic phonemic transcription.
Input
- Questioned recordings (minimum 1 recording, with at least 7 seconds of net speech).
- Suspected speaker recordings (minimum either 3 recordings with at least 7 seconds of net speech each, or 1 recording with at least 20 seconds of net speech).
- The population set (technical minimum 10 speakers, recommended minimum 50 speakers. Depending on the type of score calculation, each speaker should have one or more recordings with at least 7 seconds of net speech per recording).
- Recordings to be inspected using Deepfake Detection should contain at least 3 seconds of continuous speech for inference; best performance is achieved on recordings of 5 seconds or longer.
Supported audio formats: MS Wave or RAW with linear coding (8 or 16 bits), A-law, Mu-law; sampling frequency of 8 kHz or higher.
Output
- A scoring table with speaker comparison results in Likelihood Ratio, Log-Likelihood Ratio (decimal or natural logarithm), and Verbal Ratio.
- A table with Deepfake Detection results (Log-Likelihood Ratio).
- Graphical presentation of speaker comparison results as a Probability Density Function plot and a Tippett plot.
- The Diarization panel with labels for different speakers.
- The Phoneme Transcription panel for discovering similar sound patterns across recordings.
- The Voice Activity Detection panel with labels for speech and non-speech segments.
- The Spectrum panel and detailed spectrum layout.
- An editable report compiling all results (including the scoring table and graphs) into a single document, exportable as PDF or OpenDocument.