Invention Title:

MULTI-MODAL SYNTHETIC VOICE DETECTION SYSTEM AND METHOD

Publication number:

US20260301749

Publication date:
Section:

Physics

Class:

G10L17/24

Inventors:

Assignee:

Applicant:

Smart overview of the Invention

The patent application describes a system and method for detecting synthetic voices generated by artificial intelligence (AI) during phone calls. This system employs a combination of passive and active techniques to analyze audio inputs from callers. The passive techniques involve extracting acoustic features, spectral patterns, prosodic elements, and lexical characteristics from the audio. Meanwhile, active engagement involves generating interactive challenges through an AI agent to assess the caller's responses.

Background

The need to distinguish between human and AI-generated voices has become increasingly important with advancements in AI technology. Historically, the Turing Test was used to evaluate a machine's ability to mimic human intelligence. Today, AI systems like ChatGPT can convincingly pass such tests, particularly in text-based interactions. However, the integration of speech recognition and generation with language models has enabled AI to engage in real-time voice conversations, posing risks such as fraud and illegal robocalls.

Detection Method

The system's detection method involves receiving an audio input from a caller and analyzing it using passive detection techniques. It then actively engages the caller by presenting interactive challenges designed to probe response timing, content coherence, and complexity. The responses are processed using automatic speech recognition and natural language processing techniques. By combining results from both passive and active analyses, the system generates a composite synthetic voice detection score to determine if the voice is AI-generated.

System Components

The system comprises several components: a communications interface for receiving audio input, a processing system with a processor, and a memory storing executable instructions. Key modules include a passive detection module, an active engagement module, an automatic speech recognition module, and a natural language processing module. These modules work together to generate the composite detection score and make determinations about the nature of the caller's voice.

Applications and Network

The system can be integrated into various communication networks, including broadband, wireless, and voice access networks. It supports a range of devices from mobile phones to telephony devices, and can be implemented in different network environments such as VoIP, IP, and optical networks. The system is designed to enhance security by identifying and mitigating the risks associated with AI-generated synthetic voices in telecommunication systems.