CopunditSample
View as PDF Download PDF

Data Landscape

What labeled data exists for turning the audio of an in-person meeting into a transcript, speaker labels, a summary and action items, who holds each corpus, and on what terms it can be obtained.

Contents

Abbreviations

Abbr Stands for What it actually is (plain English)
AIMU Actionable Items for Meeting Understanding A 2016 layer of assistant-task labels added on top of 22 ICSI meetings
AMC-A AliMeeting-Action Corpus A Chinese meeting corpus whose sentences are marked as containing an action item or not
AMI Augmented Multi-party Interaction A 2005 corpus of staged meetings, recorded on headsets and on a table array at the same time
ASR Automatic Speech Recognition The software that turns recorded speech into written words
ATF Acoustic Transfer Function A measurement of how a room changes a sound between the mouth and the microphone
CC BY 4.0 Creative Commons Attribution 4.0 International A public licence allowing any use, including commercial, if the source is credited
CC BY-NC-ND 4.0 Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 Credit required, no commercial use, and no altered version may be shared
CC BY-NC-SA 4.0 Creative Commons Attribution-NonCommercial-ShareAlike 4.0 Credit required, no commercial use, and any derivative carries the same terms
CC BY-SA 4.0 Creative Commons Attribution-ShareAlike 4.0 International Credit required, commercial use allowed, and any derivative carries the same terms
CDLA-Permissive Community Data License Agreement, Permissive A data licence allowing commercial use and imposing no terms on derived work
CER Character Error Rate The word error rate's equivalent for languages written without spaces, such as Mandarin
CHiME Computational Hearing in Multisource Environments The long-running challenge series that sets the benchmarks for meeting transcription
DAMO Discovery, Adventure, Momentum and Outlook Alibaba's research academy; its speech lab released the AliMeeting and AMC-A corpora
DASR Distant Automatic Speech Recognition Speech recognition when the microphone is across the room instead of at the mouth
DER Diarization Error Rate Percentage of speaking time attributed to the wrong person, or missed, or invented
DiPCo Dinner Party Corpus A 2019 corpus of four-person dinner conversations recorded on five table arrays
DUA Data Use Agreement A signed contract setting what a recipient may and may not do with a released dataset
ELITR European Live Translator A research project whose corpus holds meeting transcripts and hand-written minutes
GDPR General Data Protection Regulation The European Union's data protection law; a recording of a person is personal data under it
GLM Global Mapping file A scoring-time list saying which alternative wordings count as the same words
IAA Inter-Annotator Agreement How often two people given the same job produce the same label
ICSI International Computer Science Institute The Berkeley institute whose 2003 corpus of real research-group meetings is still a benchmark
LDC Linguistic Data Consortium The University of Pennsylvania body that licenses and distributes speech corpora
LLM Large Language Model A text-generating model; here, the part that writes the summary and pulls out action items
LOTUSDIS Thai far-field meeting corpus, 2025 Nine separate recording devices at measured distances, all on the same Thai meetings
MAPSSWE Matched Pairs Sentence-Segment Word Error test The standard test for whether one recogniser really beats another rather than got lucky
MC Multi-Channel Audio recorded by several microphones at once, which preserves where each sound came from
MIT Massachusetts Institute of Technology Here, the permissive software licence named after it, which allows commercial reuse
MMCSG Multimodal Conversations in Smart Glasses A 2024 challenge on conversations recorded by a head-worn consumer device
NIST National Institute of Standards and Technology The United States agency whose scoring toolkit and evaluations set this field's measurement rules
NOTSOFAR Natural Office Talkers in Settings of Far-field Audio Recordings Microsoft's 2024 corpus of short office meetings on tabletop array devices
PII Personally Identifiable Information Anything in a recording that could be traced back to a named person
QMSum Query-based Meeting Summarization A benchmark of question-and-answer summaries built on top of older meeting corpora
RT60 Reverberation time to minus 60 decibels How long a sound keeps bouncing in a room before it dies away, in seconds
SC Single-Channel Audio recorded by one microphone, which throws away all directional information
SDM Single Distant Microphone The standard test condition of one far-away microphone, the closest public analogue to a phone
SUMM-RE French meeting corpus, 2024 About a hundred head-mounted-microphone French planning meetings, discourse-segmented
tcpWER time-constrained minimum-permutation Word Error Rate Word error rate that also punishes attributing words to the wrong speaker or the wrong time
UEM Un-partitioned Evaluation Map A file saying which parts of a recording count towards a score
WER Word Error Rate Percentage of words a transcript gets wrong; the standard accuracy number for speech

Executive Summary

The signal family is meeting-condition conversational audio: several people talking, often over each other, into a microphone that is not at anybody's mouth, plus whatever was annotated on top of it. The family is rich at one end and empty at the other, and the emptiness is specific rather than general. Verbatim transcription with speaker attribution is well supplied: AMI gives about 100 hours of headset-and-array meetings, NOTSOFAR-1 gives real office meetings across 30 rooms with a blind evaluation split, AliMeeting gives about 120 hours of Mandarin meetings recorded near-field and far-field at once, and CHiME-6, DiPCo, LOTUSDIS and Mixer 6 fill in dinner parties, Thai meetings and one-to-one interviews. Two things narrow that supply before any of it can be used here. Only the single-channel condition of any of them is reachable by a phone application at all, for the reason given in the primer below, and none of what is left is a clean held-out English test set: AMI and ICSI have been public for two decades and CHiME-8 lists AMI among the material a system may train on, NOTSOFAR-1's blind split has its reference transcripts withheld so nobody outside the challenge can score it, and LOTUSDIS, the one recent uncontaminated set, is Thai. Nothing in the family was recorded on a smartphone. Two independent sweeps assert that null and neither cites a catalogue behind it, but the field's own review lists 32 notable robust-recognition datasets and challenges since 2000 and none of them names a phone as a capture device; the nearest neighbours each fail for a stated reason, and the closest is a Waseda ad-hoc array of iPhones whose audio was never released. Summaries come from four corpora and action items from three, the largest of them Chinese: AMI and ICSI carry a four-part human abstractive summary in which Actions is an optional section an annotator could leave empty, and the published counts are 101 AMI meetings carrying 381 action items (derived from dialogue acts, not annotated directly), 1,506 action items over the 424 Chinese meetings of AMC-A, and 318 actionable turns over 22 ICSI meetings in the AIMU layer. No open corpus is a real business meeting of the target length with its audio attached: AMI is 15 to 45 minute role-play, ICSI is academic seminars, NOTSOFAR-1 is truncated to about six minutes, CHiME-6 and DiPCo are two-hour dinner parties, MeetingBank is city council proceedings under a NonCommercial licence, and ELITR, the one corpus of real hour-long project meetings, withheld its acoustic data for privacy. No corpus of any kind records whether a recording survived, and the reason is structural: a session that failed mid-collection was discarded before release, so no published file is paired with a cause of failure. The largest holdings of exactly the target signal are commercial and none has a published price or a named outside user: Appen, Defined.ai, Otter.ai, Gong.io, and Microsoft's withheld NOTSOFAR-1 reference transcripts.

Signals and Labels Primer

A usable record in this field is an audio recording of a real conversation plus a hand-made written reference for it, and the three things that decide whether it is worth anything are what recorded the sound, where that recorder was, and how the reference was produced.

Far-field meeting audio with verbatim transcripts and speaker attribution. The signal is sound from several talkers arriving at a microphone metres away, carrying room echo, noise and, for a fifth to nearly half of the time, two people at once. Corpora capture it with a known array (a fixed ring or line of microphones with published geometry), with a set of separate consumer-grade devices, or with a single processed channel. Only the last of those three is reachable by a product like this one, and the consequence runs through every hour count in this document: a third-party application on either mobile platform receives one processed mono stream and never the raw per-microphone channels, so an array condition supervises a signal a phone application cannot produce, and the single distant microphone (SDM) or single-channel (SC) subset is the part of any corpus here that a phone can be measured against. The label is a human-typed transcript of every word, cut into utterances and attributed to a named speaker; the governing granularity is the session, meaning one meeting as heard by one device, because that is what a system is scored on, and it is why a corpus can report both a meeting-hours figure and an all-microphones figure several times larger. Ground truth comes from a headset, lapel or in-ear microphone worn by each talker: a clean per-person recording that a transcriber listens to, and that costs a wearer's consent in every meeting.

Meeting transcripts with human summaries, minutes and action items. The signal here is text, either a human transcript or one corrected by hand from machine output, and the audio may not be released at all. The label is prose written by an annotator after reading or attending: a free-form abstract, a set of structured sections, or formal minutes. The governing granularity is the annotated meeting, far smaller than the hour count suggests, because one meeting of an hour yields one summary. For action items the granularity drops again, to the labeled sentence or turn, and it drops hard: the target class runs from half a percent to four percent of all turns depending on the corpus, 0.5 percent in AMC-A, 1.5 percent overall in AIMU and 4.2 percent in AIMU's densest group of meetings. Ground truth is the annotator's judgment, which is why this family has no accepted score and why every serious set reports an inter-annotator agreement figure beside its counts.

Simulated far-field meeting audio. The signal is built rather than recorded: clean single-speaker speech convolved with measured room responses, mixed with noise and with other talkers to a chosen overlap. Because the mixer knows exactly which words belong to whom and when, the label is exact and free at any volume, and the governing granularity is the synthesised hour. What it cannot supply is conversational behaviour, since turn-taking, interruption and the way people move around a room are not part of the recipe. The field's rule follows from that: simulated audio is accepted for TRAINING an acoustic model and rejected for EVALUATING a hardware claim.

Landscape at a Glance

Access verdicts, used in this table, in every profile header and in the coverage table: OPEN, downloadable now with no agreement and no fee; APPLICATION, behind an agreement or a committee, with a named outside group that has obtained it; PARTNER, held by named institutions with no external precedent found; COMMERCIAL, sold by a vendor; NONE, no located corpus pairs this signal with this label. The verdict is an availability column and not a permission column, and the two come apart here more often than not. A parenthesis after the verdict says what is established about the licence, and a reader should treat the qualifier as binding: a licence that was not read is not permission, and a licence two sources disagree about is not permission either. Only AMI and LOTUSDIS carry a licence permitting commercial use that anybody in this project has read.

Corpus Signal and labels N at label granularity Verdict Who else holds it
AMI Meeting Corpus Headset and 8-channel array meeting audio; verbatim transcript, speakers, dialogue acts, extractive and abstractive summaries 137 scenario meetings with a four-part summary, about 65 hours; 101 meetings carrying 381 action items; about 100 hours of audio OPEN (CC BY 4.0) Everyone; public download
NOTSOFAR-1 recorded meetings Tabletop and linear array office-meeting audio; word-aligned transcript and speaker attribution 280 meetings in the dataset paper (107 train, 36 dev, 137 eval), 315 in the challenge paper, 237 in the repository; about 28 meeting hours, about 260 across all microphones OPEN (licence in dispute) Everyone; public download
LOTUSDIS Nine separate single-channel devices at 0.12 to 10 metres on the same Thai meetings; transcript, speakers, overlap mask 90 sessions of 15 to 20 minutes, 3 speakers each, 86 speakers; about 20 meeting hours, 114 across all microphones OPEN (CC BY-SA 4.0) Everyone; public download
ICSI Meeting Corpus Far-field academic meeting audio; verbatim transcript, dialogue acts, extractive and abstractive summaries 75 meetings, about 72 hours, 3 to 10 participants; 61 meetings with an abstractive summary OPEN (licence not established) Everyone via a preprocessed public release; also catalogued by the LDC
AliMeeting 8-channel array plus per-participant headset audio of the same Mandarin meetings; verbatim transcript and speakers About 120 hours over roughly 220 to 240 sessions of 15 to 30 minutes, 2 to 4 participants, 481 speakers, 13 rooms OPEN (licence not established) Everyone via a public download
CHiME-6 Four-person dinner-party audio on 6 four-microphone Kinect arrays; verbatim transcript and speakers 20 parties, each at least 2 hours, split 16 train, 2 dev, 2 eval; 49:44 annotated hours OPEN (licence not established) Everyone via the challenge toolkit
DiPCo Four-person dinner-party audio on 5 seven-microphone circular arrays; verbatim transcript and speakers 10 sessions, 32 speakers, 5:19 hours annotated across all splits OPEN (licence not established) Everyone via the challenge toolkit
Mixer 6 Speech Two-person interviews captured by 10 heterogeneous far-field devices; verbatim transcript and speakers 450 interview portions annotated of 1,425 sessions; 20:54 fully annotated hours APPLICATION (LDC) LDC licensees; challenge participants during the challenge
AMC-A (AliMeeting-Action Corpus) Mandarin meeting transcripts with every sentence marked for containing an action item 424 meetings, 306,846 utterances, 1,506 action items, 3.55 per meeting, agreement 0.47 OPEN (licence not established) Everyone via a public repository
ELITR Minuting Corpus Meeting transcripts with hand-written minutes; audio not released 113 English and 53 Czech meetings, over 160 hours of content OPEN (licence not established) Everyone via a preprocessed public release; the audio, nobody
AIMU Assistant-executable intents labeled on the turns of 22 ICSI meetings 21,035 turns, 318 with an actionable item (1.5 percent), 10 intent types OPEN (licence not established) Everyone, if the release address still resolves
MeetingBank City-council meeting video and transcripts; official minutes, agendas and segment summaries 6,892 segment-level summarisation instances over 1,366 meetings, 3,579 hours OPEN (NonCommercial only) Everyone for research; nobody for a commercial product
NOTSOFAR-1 simulated training set Synthesised far-field meeting mixtures matched to the array geometry; exact transcripts and speaker labels by construction About 1,000 hours, built with 15,000 measured acoustic transfer functions OPEN (licence in dispute) Everyone; public download

Corpora by Signal Family

Far-field meeting audio with verbatim transcripts and speaker attribution

The best-supplied family in the field and the one every published transcription number comes from. Eight corpora carry real multi-talker audio with hand-made transcripts and speaker labels; two have a licence whose text was read, five have a working download and a licence claim nobody has verified against the holder's own page, and one sits behind a consortium. AMI Meeting Corpus, NOTSOFAR-1 recorded meetings, LOTUSDIS, ICSI Meeting Corpus, AliMeeting, CHiME-6, DiPCo, Mixer 6 Speech.

AMI Meeting Corpus | OPEN (CC BY 4.0)

NOTSOFAR-1 recorded meetings | OPEN (licence in dispute)

LOTUSDIS | OPEN (CC BY-SA 4.0)

ICSI Meeting Corpus | OPEN (licence not established)

AliMeeting | OPEN (licence not established)

CHiME-6 | OPEN (licence not established)

DiPCo | OPEN (licence not established)

Mixer 6 Speech | APPLICATION (LDC)

Meeting transcripts with human summaries, minutes and action items

The thin family, and the one that decides whether the second and third promised outputs can be measured at all. Five entries carry a written label made by a person after the meeting; the largest forbids commercial use, the richest for action items is Chinese, the one with real hour-long project meetings has no audio, and the two English corpora that pair audio with summaries are profiled above with the audio they came from. AMC-A (AliMeeting-Action Corpus), ELITR Minuting Corpus, AIMU, MeetingBank.

AMC-A (AliMeeting-Action Corpus) | OPEN (licence not established)

ELITR Minuting Corpus | OPEN (licence not established)

AIMU | OPEN (licence not established)

MeetingBank | OPEN (NonCommercial only)

Simulated far-field meeting audio

One entry, and it has its own family because its labels are exact and its conversations are not real. It is how the current front ends were trained when nobody could record enough meetings. NOTSOFAR-1 simulated training set.

NOTSOFAR-1 simulated training set | OPEN (licence in dispute)

Holders of Closed Cohorts

The data in this field that never left the building, grouped by what it holds. The academic side publishes its corpora, so this section is mostly the industrial side, and the industrial side is where the volume is.

Smartphone-captured conversational speech. Ochi and colleagues at Waseda University published in 2016 on multi-talker recognition over "an ad hoc microphone array, which consists of smartphones... realized using iPhone and Dropbox". The method for synchronising several consumer handsets is therefore on the record; the audio behind it was never released as a public benchmark, and the paper was not obtainable in this pass, so the quotation above is all that is established. Appen holds a proprietary corpus titled "English (United States) Conversational Smartphone Speech" containing 1,000 hours of fully transcribed data, available only by commercial procurement (quoted from appen.com/speech-and-audio-training-data; the page itself was not read here). Nothing states its speaker count per recording, and telephony collections of this kind are historically two-party.

Real business meetings with the customer's own feedback attached. Otter.ai states that its models are trained on millions of hours of audio recordings gathered from its user base, and Gong.io captures and transcribes thousands of hours of business-to-business sales calls and internal meetings through what it calls its Revenue Graph. Both are real meetings of natural length captured on consumer microphones, both are locked behind business privacy and compliance commitments, and no outside researcher or competing developer has obtained either (quoted from the vendors' own pages and a third-party profile; no data-availability statement of any kind exists for either). Defined.ai builds and sells bespoke conversational datasets to order, including call-centre audio and a 225-hour annotated human-demonstration set, on commercial purchase agreements with no published price.

The withheld halves of published corpora. Microsoft holds the reference transcripts for the NOTSOFAR-1 evaluation split, deliberately, so that a challenge score cannot be overfitted, and it holds the 15,000 measured acoustic transfer functions behind the simulated set; the synthesised audio is released and no source read here says whether the measurements are. The ELITR project holds the audio behind its 113 English and 53 Czech meetings, withheld for privacy (Rennard 2023), and that is the sharpest single fact in this document: the right meetings were recorded, and privacy stopped their release. The CHiME-6 organisers hold the video recorded alongside the audio, used to re-synchronise the arrays and described as available only to the organisers (Cornell 2025). The Linguistic Data Consortium holds the unannotated remainder of Mixer 6, 975 of 1,425 sessions with only partial subject-only annotation, and the corpus's conversational portions were transcribed by the challenge organisers rather than by the holder.

Meeting audio with summaries. There is no located academic closed cohort here, and there is an explicit statement of why the public ones are so few: "the cost of producing such corpora, together with concerns about the privacy of meeting content, mean that there are very few such data sets available", after which the survey names three for English and no more (Rennard 2023). A second survey gives the same reason from the other side, that most meetings performed in industry are proprietary (arXiv 2212.08206).

Making Data That Does Not Exist

Four signal-label pairs the located corpora do not cover, and what producing each one consists of in this field, as facts about how the field has done it before.

Smartphone-captured meeting audio with verbatim transcripts and speaker attribution. The field's protocol for this is a parallel-capture study, and its parts are all documented. The topology is the devices under comparison placed adjacently at the same point on the table, with a close-talking lapel or headset microphone on every participant to produce the reference. The reference transcript is made by humans listening to the per-speaker channels, and machine pre-transcription is refused rather than merely discouraged, because annotators accept plausible machine guesses in noisy segments and thereby import the model biases the corpus exists to measure (Vinnikov 2024). Scoring conventions matter as much as the audio: the NIST scoring toolkit's Global Mapping (GLM) files decide which alternative wordings count as the same words, and a mapping that expands contractions in both reference and hypothesis double-counts errors. Because independent devices share no word clock, their sample rates drift over a 30-minute meeting, so the CHiME-6 baseline aligned array signals to the reference headsets by cross-correlation with sox, and LOTUSDIS synchronised nine devices by slate pulse verified to sub-sample alignment. The metric is a speaker-attributed word error rate, cpWER or tcpWER, and the significance test is the Matched Pairs Sentence-Segment Word Error (MAPSSWE) test in the NIST toolkit, with McNemar or Wilcoxon signed-rank as non-parametric cross-checks. The shape of a minimum credible study, as the field's own challenge splits imply it: at least 20 sessions of at least 15 minutes, at least 20 speakers in groups of 3 to 5, rotating through at least 4 rooms of different reverberation time, 10 to 15 hours of test audio, one recogniser decoding both device streams so the microphone is the only variable. The pass criterion is symmetric and stated in advance: the phone passes if its cpWER is significantly lower than the comparison device's, or if MAPSSWE returns no significant difference; it fails if the comparison device is significantly lower at 95 percent confidence. Room diversity matters more than duration: NOTSOFAR-1 used 30 rooms and short sessions, AliMeeting 13 rooms and long ones.

The simulation shortcut is closed in both directions and this is not a matter of preference. Simulated far-field audio is accepted by the field for TRAINING an acoustic model and rejected universally for EVALUATING a hardware claim, because linear summation of convolved sources does not reproduce the Lombard effect, non-linear pre-amplifier compression or clipping when two loud voices hit the converter at once. And building a simulator for a new device is not cheaper than recording: NOTSOFAR-1's simulator rests on 15,000 acoustic transfer functions physically measured on the target hardware in real rooms, which is the recording session the simulation was supposed to avoid.

The other route in this field is not a study at all, and it is the one every commercial holder named above took. A product ships, a training-data opt-in sits in its terms of service, and the users' own meetings are retained under it together with the users' corrections to the transcript, the summary and the action items, which are the only labels for the second and third outputs that ever reach volume. It is how Otter.ai and Gong.io came to hold what they hold. Two properties make it a different instrument from the parallel-capture study rather than a cheaper version of it. It produces nothing before a product ships, so it cannot settle a capability question in advance. And its cost is legal rather than logistical: a voice recording is personal data, a speaker label makes it biometric, the consent a terms-of-service checkbox obtains binds the user and not the other people in the room, and the ethics guidance read here holds that presenting commercial development as academic research invalidates the lawful basis of the consent.

Action items as an independently counted target. Two production routes exist and they cost differently. The AMI and ICSI route makes action items a section of a summary: one annotator writes up to 200 words under Actions and may leave it empty, which yields the 101 meetings and 381 items the literature has been able to derive, by treating dialogue acts linked to that section as positive examples. The AMC-A route makes them a label in their own right: sentence-level binary classification, three independent annotators on candidate sentences pre-highlighted for temporal expressions and action verbs, an expert adjudicating the majority vote, which yields 1,506 items over 424 meetings at a pairwise Cohen's kappa of 0.47. The pre-highlighting step is the documented cost lever and the agreement figure is the documented ceiling: the same paper reports a kappa of 0.36 for the earlier ICSI attempt and calls action items inherently subjective. On the summary side, the most rigorous published human protocol is Atomic Content Units, in which a reference summary is decomposed into binary atomic claims and each is checked against the system output; the benchmark that established it required over 150 hours of in-house annotation to produce 22,000 summary-level annotations, on news and chat data rather than meetings (RoSE, arXiv 2212.07981). One further external figure circulates, that multi-party annotation regularly exceeds 50 hours of labour per hour of audio, and no source read here carries a citation for it.

Long-session capture reliability. No corpus in any family records whether a recording completed, and the reason is structural rather than accidental: acoustic corpora ship cleanly truncated audio files, and a session that failed mid-collection was discarded before publication to keep the benchmark intact, so no release pairs a truncated file with a cause such as an application crash, a thermal shutdown or a microphone seized by an incoming call. This one cannot be produced from archives at all; the field's protocol for it is instrumentation, not annotation. The documented schema logs session start and end, buffer overrun and underrun counts, operating-system audio-session interruptions and watchdog-timer expiries, as structured records that never carry the audio itself, which keeps the telemetry out of biometric and personal-data scope. The shape of a minimum credible study, again from the field's own practice: several hundred real sessions from several dozen users on their own handsets across both platforms, with a defect defined as any session losing audio frames to a buffer underrun, an unhandled interruption or a watchdog timeout, and a defect-free completion rate above 99.0 percent as the pass criterion.

Buyer-stated reasons. What a recorder buyer is paying for is answered by interview or survey data about people, not by a corpus of audio, and none was located in any form. What exists is technology-reviewer opinion, which is a different instrument.

What consent and ethics look like when this data is made now. AMI and ICSI were collected before the General Data Protection Regulation (GDPR) was enforced; AMI went through the European Commission's ethics review procedure with a technical annex checklist rather than a United States institutional review board. Under current practice a voice recording is personal data and, where it is used to identify a person, special-category biometric data requiring explicit consent. Guidance read here separates the participant's informed consent from the processing privacy notice, treats true anonymisation of voice as impossible so that data is pseudonymised instead, and requires that participants be told their voice itself is an identifier; broad consent covers unknown future algorithmic uses and dynamic consent covers new phases. Two points bear directly on any collection done to settle a commercial claim: the ethics submission and the privacy notice must state the commercial intent, because presenting commercial development as purely academic research invalidates the lawful basis of the consent; and participants retain a right to withdraw up to publication or aggregation. Release practice has moved with it, from an open server to a Data Use Agreement (DUA) plus gated hosting that requires registration and contractually forbids re-identification, which is how NOTSOFAR-1 is distributed.

Dead Ends: Corpora That Cannot Be Used

A licence that forbids commercial use, with no located waiver. MeetingBank is released under CC BY-NC-SA 4.0; the NonCommercial clause bars use in a product that is sold and the ShareAlike clause would bind any derivative to the same terms. It remains usable for evaluation and for a paper. No route to a commercial licence was found and the data card names no rights holder to ask. The same licence string is attributed to ELITR by the corpus sweep, unverified, and if it holds the same bar applies there.

A ShareAlike clause that would follow the work out. If the CC BY-SA 4.0 attributions read here are right, CHiME-6, AliMeeting, AISHELL-4 and LOTUSDIS all permit commercial use and all require that a derivative dataset be released on the same terms. That is not a bar to using them and it is a bar to keeping quiet about what was built from them, which is a different decision and has to be made deliberately. None of the four was read from its holder's own licence file.

The audio was never released. The ELITR Minuting Corpus distributes transcripts and minutes only, with sections censored for privacy, so nothing acoustic can be trained or measured on it however good the minutes are. Route back: none found; the decision is described as a privacy one, and it is the same force that would apply to anything collected fresh.

The annotation went offline. The action-item annotations added to 18 ICSI meetings by Purver and colleagues, reported at a Cohen's kappa of 0.36, "are no longer publicly available" (Liu 2023, section 2.2). That is the earliest English action-item label set in this field and it cannot be obtained. AIMU's release address should be checked against the same risk before anything is built on it.

The access window closed. Free access to Mixer 6 was provided by the Linguistic Data Consortium to challenge participants for the duration of the challenge. Route back: the consortium's standard catalogue licensing, at the fee quoted in the profile above.

Public long enough that a held-out claim cannot be made. The CHiME-8 organisers write that in the previous edition the evaluation data was "not really blind", because only Mixer 6 was partially blind and that data had been public for more than a decade. This does not stop anyone using AMI, ICSI, CHiME-6, DiPCo or Mixer 6; it stops a claim that a number computed on them is a clean held-out result. NOTSOFAR-1's blind evaluation split is the located exception.

The wrong regime entirely. Read-speech and telephone corpora appear throughout this field's history and are allowed as external training material in the challenges, but a number computed on them is not meeting evidence. Corpora built by replaying clean speech through loudspeakers into a room, such as LibriCSS, sit in the same place: there is no natural overlap, the mixing is scripted, and the recordings carry no Lombard effect, which makes them useful for separation research and not a measurement of real conversation. The archive types considered and rejected for a device comparison are on the record with their reasons: podcasts (near-field dynamic microphones), video-call recordings (vendor noise suppression and automatic gain control irreversibly mask the raw microphone), body-camera audio (moving target, restricted), and parliamentary or council proceedings (push-to-talk gooseneck microphones, which misrepresent a single device on a table).

Coverage of the Pitch

One row per claim whose test needs data. Descriptive only: it records what exists, not which claim to keep.

Claim Signal and label needed Verdict Nearest corpora
C1, an app on a phone can capture a meeting well enough for a usable transcript, summary and action items Smartphone-captured in-person meeting audio with verbatim transcripts, speaker attribution and a summary NONE AMI supplies the meeting condition and all three label types, but on a headset and a table array; NOTSOFAR-1 and AliMeeting supply the acoustics and no summaries; no located corpus of any kind was recorded on a phone
C2, the phone's own microphones are better than the dedicated recorder's for this job The same meetings captured simultaneously by a phone and by the comparison device, scored on one reference transcript NONE LOTUSDIS is the closest design, nine separate devices at 0.12 to 10 metres on one conversation, and contains no phone and no English; Mixer 6 puts 10 device types on a two-person interview; AliMeeting pairs near-field and far-field of the same meetings
C3, the three outputs are finished within the few minutes after the meeting ends Wall-clock production time from end of recording to finished outputs, with recording length NONE No corpus records a production time; NOTSOFAR-1's blind evaluation audio is material a latency harness could be run on, and supplies no timing itself
C4, the phone keeps recording to the end of a real meeting Long capture sessions with completion or loss recorded as the outcome NONE Nothing located records capture outcomes, and the reason is structural: failed sessions are discarded before release; CHiME-6's two-hour-plus parties are the longest real sessions located, on dedicated arrays
C5, the person buying a 159 dollar recorder would take the app instead Stated purchase reasons from buyers of dedicated recorders NONE No corpus of any kind; the located material is reviewer opinion, not buyer research

What the Corpora Are Labeled For

The inventory, in no particular order. Each line is a signal-label pair the located corpora cover at a count that supports work, whether or not the pitch asked for it.

Also Found, Not Profiled

Every corpus the sweep located that did not earn a profile, alphabetical, carrying what was already established.

Corpus Holder What it holds What we know
AISHELL-4 AiShell Mandarin office meetings on an 8-channel circular table array, with video for audio-visual diarization 120 hours, 211 sessions, 4 to 8 participants, 61 speakers, 0.6 to 6 metres, headset plus circular array, CC BY-SA 4.0 per a corpus-comparison table (Tipaksorn 2025, Table 1); listed as real-world long-form far-field in the CHiME review's dataset table
CHiME-5 CHiME challenge series The same dinner-party recordings as CHiME-6 Shares the exact same data as CHiME-6, before the inter-array resynchronisation (Cornell 2025)
CID (Corpus of Interactional Data) Aix-Marseille University Eight one-hour French dialogues between friends, discourse-segmented Eight hours total, two-party, not meetings; cited as the prior French resource SUMM-RE supersedes (Prevot 2025)
Ego4D Ego4D consortium Egocentric video and audio from head-worn devices Named as a 2022 real-world, long-form, multi-speaker, far-field, multi-domain set in the CHiME review's dataset table; no card read
kiransarv action-item dataset GitHub, individual 2,750 statements labeled for whether they contain an action item Used as the action-item classifier's training data in an AMI summarisation study (arXiv 2312.17581); no holder, licence or provenance established
Kirstein meeting-summary error annotations University of Goettingen 175 machine-generated summaries of 35 QMSum meetings, each labeled by a human annotator for eight error types including hallucination, wrong references and missing information 35 general-summary samples from the QMSum test set summarised by five systems; Krippendorff's alpha 0.76 to 0.83 for error detection; the paper states the annotations, the annotator guidelines and the code are released at github.com/FKIRSTE/emnlp2024-Meeting-Sum-Metrics (arXiv 2404.11124, read in full and indexed in LITERATURE.md); no licence established and the repository was not read
LibriCSS Microsoft Clean read speech replayed through loudspeakers into a conference room to create controlled overlap 10 hours, 10 sessions, 8 speakers per session, 40 speakers, 0.3 to 4 metres, 7-channel circular array, CC BY 4.0 per Tipaksorn 2025 Table 1; no natural overlap and no Lombard effect, so evaluation material for separation only
LibriSpeech OpenSLR Read audiobook speech Permitted external training data in CHiME-7 and 8; read speech, so not meeting evidence
MISP MISP challenge organisers Far-field conversation with several sensing modalities Named as a 2022 real-world multi-domain challenge set in the CHiME review's dataset table; no card read
QMSum Yale LILY Query-focused summaries built on top of AMI, ICSI and parliamentary committee meetings 1,808 query-summary pairs over 232 meetings: 137 AMI, 59 ICSI, 36 committee; average meeting 9,069.8 tokens, average summary 69.6 tokens (arXiv 2212.08206, arXiv 2404.11124). The corpus sweep states an MIT licence and cites an ACL events index for it, which does not support the claim; no licence established
Santa Barbara corpus University of California, Santa Barbara Recorded American English conversation The earliest entry in the CHiME review's dataset table, 2000, marked real-world, long-form, multi-speaker, far-field and multi-domain
SUMM-RE LINAGORA Labs About 100 sessions of three 20-minute French event-planning meetings, 2 to 4 participants, one task per session including delegating the practical work Most sessions recorded face to face on head-mounted microphones, a few over Zoom; the whole corpus automatically transcribed, with 73 meetings and about 24 hours manually corrected and discourse-annotated; at huggingface.co/datasets/linagora/SUMM-RE (Prevot 2025). Head-mounted capture means it is near-field, not a far-field resource
VoxCeleb 1 and 2 University of Oxford Interview audio used to train speaker-discriminative models Permitted external training data in CHiME-7 and 8, used for speaker-identity embeddings
WeCanTalk Linguistic Data Consortium Cantonese, Mandarin and English telephone conversations plus self-recorded video from 202 bilingual speakers in Hong Kong At least 10 telephone calls of 8 to 10 minutes and at least 3 videos per speaker, collected June to October 2020 for the NIST 2021 Speaker Recognition Evaluation, to be published in the LDC catalogue; the corpus paper is itself under CC BY-NC 4.0. Two-party telephony and selfie video, not multi-party in-person meetings, and the smartphone reference in it is about handset penetration in Hong Kong, not about the capture device

Watchlist

Dated 2026-09-01. A fired trigger means this document is owed a refresh.

What to watch Corpus What it would change Where it shows up
A challenge edition or corpus release whose capture devices include a phone on the table alongside an array NOTSOFAR-1 or a successor in the CHiME series It would create the first public phone-versus-device comparison and turn the microphone question from unanswerable into measured The CHiME challenge task and data pages
Which of CC BY 4.0 and CC BY-NC-ND 4.0 actually governs NOTSOFAR-1 recorded and simulated sets The difference is between the field's current benchmark being usable in a commercial product and being unusable in one The dataset's own licence file on the repository and the Hugging Face card
The licence and access terms of the smart-glasses conversation set CHiME-8 MMCSG It is the nearest thing to a consumer wearable corpus; its terms decide whether a wearable-capture baseline can be built at all The CHiME-8 Task 3 data page
Whether the release address still resolves AIMU It is the only located English turn-level action-item annotation, and its sibling ICSI annotation has already gone offline research.microsoft.com/projects/meetingunderstanding/ and any successor Microsoft Research page
An English counterpart, or an English annotation layer built the same way AMC-A It would move action items from a derived label with 381 examples to a directly annotated target with four figures of examples ModelScope, and the meeting-understanding challenge tracks that use it
A second corpus adopting the measured-distance multi-device design LOTUSDIS It is the only located design that scores the same conversation across separate consumer-grade devices at stated distances, which is the shape a phone comparison needs The corpus repository and the far-field recognition literature that cites it
Access terms outside a challenge window Mixer 6 Speech It is the only located English corpus with ten different capture devices on one conversation, so its terms decide whether device comparison is possible on real data The Linguistic Data Consortium catalogue entry
A commercial-use variant or a licence revision MeetingBank It would move the largest located summarisation corpus from evaluation-only to something a product can be built on Its Hugging Face data card
A release of the measured acoustic transfer functions behind the simulated set NOTSOFAR-1 simulated training set It would let anyone synthesise matched training audio for a different device geometry, including a phone's The NOTSOFAR-1 repository
A reconciliation of 237, 280 and 315 NOTSOFAR-1 recorded meetings Every per-meeting number this field quotes on that benchmark, including the overlap ratios and the hours, rests on which count is right The repository contents against the dataset paper's own table
View as PDF Download PDF