- Name
- Primary purpose
- Used Software
- Attributes
- Response
- Create an ethically sourced flagship dataset to enable AI research on voice as a biomarker of health and support clinically meaningful insights.
| Grantor | Grant Name | Grant Number |
|---|---|---|
| - | - | - |
Datasheet for Dataset - Human Readable Format
Why was the dataset created?
| Grantor | Grant Name | Grant Number |
|---|---|---|
| - | - | - |
What do the instances represent?
How was the data acquired?
| ID | Name | URL | Version |
|---|---|---|---|
| openSMILE | openSMILE | https://audeering.github.io/opensmile/ | |
| Parselmouth | Parselmouth | https://github.com/YannickJadoul/Parselmouth | |
| Praat | Praat | http://www.praat.org/ | |
| Whisper-Large | OpenAI Whisper Large | https://github.com/openai/whisper | |
| torchaudio | TorchAudio | https://pytorch.org/audio | 2.1 |
| librosa | librosa | https://librosa.org | |
| b2aiprep | b2aiprep | https://github.com/sensein/b2aiprep |
| Description | ID | Media Type | Name | Path | Title |
|---|---|---|---|---|---|
| Parquet dataset with spectrograms (513×N) and identifiers (participant_id, session_id, task_name). | spectrograms.parquet | application/x-parquet | spectrograms.parquet | spectrograms.parquet | Derived spectrograms |
| Tab-delimited table of demographics, acoustic confounders, and validated questionnaire responses (on... | phenotype.tsv | text/tab-separated-values | phenotype.tsv | phenotype.tsv | Participant-level phenotype data |
| JSON data dictionary describing columns in phenotype.tsv. | phenotype.json | application/json | phenotype.json | phenotype.json | Phenotype data dictionary |
| Tab-delimited table of features derived from raw audio (one row per recording). | static_features.tsv | text/tab-separated-values | static_features.tsv | static_features.tsv | Recording-level static features |
| JSON data dictionary describing columns in static_features.tsv. | static_features.json | application/json | static_features.json | static_features.json | Static features data dictionary |
What (other) tasks could the dataset be used for?
How will the dataset be distributed?
How will the dataset be maintained?