How AI Learns to Listen for Disease Detection

Summary: An international consortium has produced the first consensus-based framework and standardized taxonomy for vocal biomarkers, establishing clear definitions and a hierarchical model to advance clinical, regulatory, and research applications of voice-based health technologies.

A multi-stage Delphi process convened 24 experts from Europe and North America through the eVoiceNet and NIH Bridge2AI-Voice consortia to resolve longstanding terminology confusion. The framework clarifies distinctions among voice, speech, and respiratory acoustic signals, and formally separates unvalidated vocal measures from clinically validated vocal biomarkers.

This structured model creates a common scientific vocabulary across physiological, cognitive, acoustic, and computational domains. It aims to accelerate clinical validation, regulatory approval, and deployment in digital health for conditions such as Parkinson’s disease, Alzheimer’s disease, depression, heart failure, and type 2 diabetes.

Key Facts

  • First Consensus Framework: Introduces the inaugural standardized terminology and classification model for health technologies that use human voice, speech, and respiratory acoustic signals.
  • Consensus Process: Developed through a multi-stage Delphi methodology (2024–2025) that engaged 24 international experts in clinical medicine, speech-language pathology, acoustic engineering, data science, and regulatory affairs.
  • Conceptual Distinction: Differentiates raw or processing-derived “vocal measures” (e.g., fundamental frequency, jitter, pause duration) from validated “vocal biomarkers” (features shown to reliably indicate a specific health condition or physiological state).
  • Diagnostic Breadth: The taxonomy addresses detection and monitoring across neurological (Parkinson’s, Alzheimer’s), psychiatric (major depressive disorder), cardiovascular (heart failure), and metabolic (type 2 diabetes) domains.
  • Regulatory and Clinical Translation: Designed to support regulatory pathways and clinical validation by providing clear ontologies, classification standards, and operational guidance for voice-based medical tools.

Source: USF

A person’s voice conveys more than words alone. Research increasingly shows that subtle changes in speech, breathing, and vocal quality can reveal information about a wide range of health conditions, from neurodegenerative diseases to psychiatric, cardiovascular, and metabolic disorders.

These voice-derived indicators—vocal biomarkers—are poised to become tools for disease screening, diagnosis, and longitudinal monitoring.

To unlock this potential, teams at the Department of Precision Health (DoPH) at the Luxembourg Institute of Health (LIH) and the University of South Florida Morsani College of Medicine led an international effort to create consensus definitions and a classification system for vocal biomarkers.

Published in Digital Biomarkers under the VOCAL (Vocal Biomarker Guidelines for Ontology, Classification, Application and Logistics) initiative, the study gathered 24 experts from Europe and North America to address a major barrier: inconsistent terminology that hinders reproducibility, comparison across studies, and regulatory review.

Rapid growth in vocal biomarker research had produced overlapping and interchangeable use of terms like “voice,” “speech,” and “vocal” biomarkers, despite differences in the underlying physiological and cognitive processes. The VOCAL initiative, coordinated by eVoiceNet (Europe) and Bridge2AI-Voice (North America), used a rigorous multi-stage consensus process in 2024–2025 to create a shared taxonomy and definitions.

The resulting framework distinguishes between generic vocal measures—quantifiable acoustic, linguistic, or respiratory parameters—and validated vocal biomarkers—measures or combinations of measures that have undergone clinical validation and reliably indicate a defined physiological state, biological process, or diagnosis. It also introduces a hierarchical continuum spanning broad biomarker concepts to domain-specific measures covering cardiorespiratory acoustics, voice production, articulation and speech, and cognitive/language features.

This common vocabulary is intended to improve collaboration among clinicians, speech-language specialists, engineers, data scientists, regulators, and industry partners. It also supports development of validation pathways, reporting standards, and regulatory guidance for voice-based digital health tools.

“Voice contains rich health information, but without a shared language the field cannot progress efficiently,” said Dr. Guy Fagherazzi, head of DoPH at LIH and eVoiceNet chair. “Defining our terms lays the groundwork for rigorous research, transparency, and technologies that can benefit patients.”

Dr. Yael Bensoussan, associate professor of Otolaryngology at USF Health and co-head of Bridge2AI-Voice, added that vocal biomarkers capture signals from multiple physiological and cognitive systems simultaneously. The framework preserves that complexity while enabling clearer communication and reproducible methods.

This publication represents the first phase of the broader VOCAL initiative, which will continue to develop international guidelines and standards to accelerate translation of voice-based technologies from research into clinical practice.

Key Questions Answered:

Q: What is the main difference between a “vocal measure” and a “vocal biomarker” under the VOCAL framework?

A: A vocal measure is any quantifiable acoustic, linguistic, or respiratory parameter derived from voice recordings (for example, pitch variability or pause length). A vocal biomarker is a vocal measure—or a validated set of measures—that has undergone clinical evaluation and demonstrates reliable association with a specific health condition or physiological state.

Q: Why was standardized terminology necessary for voice-based digital health technology?

A: The field’s rapid expansion produced inconsistent and interchangeable use of key terms. Because voice production involves overlapping systems (laryngeal, respiratory, neurological, cognitive), a unified taxonomy enables better collaboration, reproducibility, and regulatory review, facilitating clinical adoption.

Q: Which conditions can potentially be monitored using vocal biomarkers?

A: Current research indicates potential utility for tracking neurodegenerative disorders (Parkinson’s, Alzheimer’s), psychiatric conditions (depression, anxiety), cardiovascular issues (heart failure, pulmonary congestion), and metabolic diseases (type 2 diabetes), among others.

Editorial Notes:

  • This article was edited by a Neuroscience News editor.
  • The journal article was reviewed in full by the editorial team.
  • Additional explanatory context was added by staff to aid clarity and reader understanding.

About this AI research news

Author: Cody Hawley
Source: USF
Contact: Cody Hawley – USF
Image: The image is credited to Neuroscience News

Original Research: Open access.
“Consensus-Based Definitions for Vocal Biomarkers: The International VOCAL Initiative” by Mégane Pizzimenti et al., published in Digital Biomarkers.
DOI: 10.1159/000553327


Abstract

Consensus-Based Definitions for Vocal Biomarkers: The International VOCAL Initiative

Introduction: Voice-based health technologies are expanding rapidly but lack standardized terminology, which limits interdisciplinary collaboration, research quality, and clinical translation. This work—part of the VOCAL initiative—develops consensus definitions and a classification framework to provide standards and guidance for the field.

Methods: VOCAL used a rigorous international multi-stage consensus process in 2024–2025, bringing together 24 experts from the Bridge2AI-Voice Consortium and eVoiceNet. The iterative process included five review rounds and an in-person workshop at the 2025 Bridge2AI Voice Symposium, integrating perspectives from medicine, speech science, signal processing, statistics, regulation, and ethics.

Results: The initiative produced consensus definitions and a hierarchical continuum model for vocal biomarkers. It distinguishes vocal measures from validated vocal biomarkers and defines a multi-level taxonomy from general biomarker concepts to domain-specific measures, including cardiorespiratory acoustics, voice production, articulatory/speech features, and cognitive/language metrics with linguistic and paralinguistic subtypes.

Conclusion: The VOCAL framework delivers a shared vocabulary essential for interdisciplinary communication, higher-quality research, and the ethical, reliable deployment of voice-based health technologies. It establishes foundational standards to support future validation pathways, regulatory guidance, and broader clinical adoption of vocal biomarkers.