Skip to main content

Key Features

Phonexia Voice Inspector offers several features to effectively support voice forensic experts:

  • Standalone application with a comprehensive, easy-to-use Graphical User Interface (GUI)
  • Automatic speaker comparison of one or more questioned recordings (unknown-speaker recordings or voiceprints) against one or more suspected reference speakers, supporting 1:1, 1:N, and N:M comparison modes.
  • Integrated Deepfake Detection to identify synthetic audio alongside speaker comparison results.
  • Integrated speech technologies: Speaker Identification, Speaker Diarization, Phoneme Recognition, Voice Activity Detection, and Speech Quality Estimation.
  • Ability to search for repetitive sound patterns across all recordings using automatic phonemic transcription.

Input

  • Questioned recordings (minimum 1 recording, with at least 7 seconds of net speech).
  • Suspected speaker recordings (minimum either 3 recordings with at least 7 seconds of net speech each, or 1 recording with at least 20 seconds of net speech).
  • The population set (technical minimum 10 speakers, recommended minimum 50 speakers. Depending on the type of score calculation, each speaker should have one or more recordings with at least 7 seconds of net speech per recording).
  • Recordings to be inspected using Deepfake Detection should contain at least 3 seconds of continuous speech for inference; best performance is achieved on recordings of 5 seconds or longer.

Supported audio formats: MS Wave or RAW with linear coding (8 or 16 bits), A-law, Mu-law; sampling frequency of 8 kHz or higher.

Output

  • A scoring table with speaker comparison results in Likelihood Ratio, Log-Likelihood Ratio (decimal or natural logarithm), and Verbal Ratio.
  • A table with Deepfake Detection results (Log-Likelihood Ratio).
  • Graphical presentation of speaker comparison results as a Probability Density Function plot and a Tippett plot.
  • The Diarization panel with labels for different speakers.
  • The Phoneme Transcription panel for discovering similar sound patterns across recordings.
  • The Voice Activity Detection panel with labels for speech and non-speech segments.
  • The Spectrum panel and detailed spectrum layout.
  • An editable report compiling all results (including the scoring table and graphs) into a single document, exportable as PDF or OpenDocument.