CopunditSample
View as PDF Download PDF

Literature Index - far-field meeting transcription and meeting summarisation

This vault answers what a published error rate for meeting audio actually is, under which recording condition it was measured, and what can and cannot be measured about a summary or an extracted action item. The papers cover the CHiME-6, CHiME-7, CHiME-8 DASR and CHiME-8 NOTSOFAR-1 distant-speech challenges and several of their submitted systems, the 2026 state of the art in single-microphone speaker-attributed transcription, six audio and annotation corpora with their recording conditions and licences, surveys and methods for meeting summarisation and action-item extraction, five studies of whether a summary can be scored at all, and one platform vendor's report on the language model that runs on a phone. A paper is in this vault only because its PDF is in literature/ and was read end to end.

Papers indexed: 30

Contents

How to use this index

Every number in this index is written with the audio condition it was measured under and the named test set it was measured on, because a word error rate from a seven-microphone tabletop array and a word error rate from a single distant microphone stream are different numbers about different products, and the project's contract forbids quoting one for the other. Blocks headed "Quotable stats" are complete sentences ready to paste into a brief; blocks headed "Direct quotes" are verbatim ground truth from the PDF, page-referenced, and are what to fall back on when a downstream document has drifted. Before citing any paper here for a mechanism claim, read its "Mechanism family" tag and its "What this paper does NOT establish" block: several of these papers are routinely mis-cited for claims they explicitly did not test.

Papers are grouped into six thematic clusters:

Admission gate and evidence tiers

Adopted 2026-09-01 from the studio default template, with the domain venue list filled for this project. It is project-owned from here: edit or delete it and the index will follow the edited version.

Two papers in this vault carry a flag a reader needs to see before quoting them. Niu 2024 is a METHOD paper that names no repository for its own system, so it fails the implementation arm despite a strong venue and a strong citation rate. Kumar 2022 is an arXiv preprint with no conference or journal venue and a citation rate below the floor, and its central artifact is a leaderboard compiled from other papers' reported numbers rather than from its own measurements.

Mechanism-family taxonomy

The controlled vocabulary used in this index. Two papers with different tags must never be cited interchangeably for the same claim. The resolution here is set by the project's own exposure: the charter proposes capturing an in-person meeting with a phone's own microphones, so the tags separate what was measured on one microphone stream from what was measured on a microphone array, and separate scoring the words from scoring the summary.

Common confusions in this domain, to read before pasting any of these papers into a pitch:

Quick-reference table

Short cite Venue / Year Mechanism family N Headline result
Cornell 2025 Computer Speech and Language, 2025 Far-field multi-channel transcription benchmark, Joint speaker-attributed transcription metric, Summarisation evaluation metric 32 systems, 9 teams, 4 scenarios Best macro tcpWER 33.6 percent across four scenarios; CHiME-6 stays above 30 percent for every system; summary quality correlates with tcpWER at only PCC -0.51 (G-Eval overall)
Cornell 2024 CHiME 2024 Workshop Far-field multi-channel transcription benchmark, Corpus resource 4 scenarios, 2 baselines Baseline macro tcpWER 56.5 percent (NeMo) and 62.6 percent (ESPnet) on eval; speaker counting named as the dominant error source
Abramovski 2025 Computer Speech and Language, 2025 Far-field single-channel transcription benchmark, Far-field multi-channel transcription benchmark 315 meetings, 30 rooms, 14 submissions Best single-channel tcpWER 22.2 percent vs best multi-channel 10.8 percent on the same 170-meeting eval set
Niu 2024 CHiME 2024 Workshop Far-field single-channel transcription benchmark, Array front-end signal processing NOTSOFAR-1 eval, 170 meetings Won both tracks: tcpWER 22.2 percent single-channel, 10.8 percent multi-channel, using an ensemble of three modified Whisper models
Shi 2023 APSIPA ASC 2023 Far-field multi-channel transcription benchmark, Joint speaker-attributed transcription metric AliMeeting, 104.75 h train / 4 h eval / 10 h test Best average SD-CER 28.3 percent with an 8-channel array vs 34.4 percent from the beamformed single channel
Rennard 2023 TACL, 2023 Abstractive summarisation method, Corpus resource 3 English meeting corpora, about 280 h total Only three English meeting-summarisation corpora exist, all with human-produced or human-corrected transcripts; best AMI ROUGE-1 55.27
Kumar 2022 arXiv preprint, 2022 Abstractive summarisation method, Extractive summarisation method over 40 papers surveyed Leaderboard compiled from published results: best reported AMI ROUGE-1 56.26, best reported ICSI ROUGE-1 60.7
Kirstein 2024 Findings of EMNLP 2024 Summarisation evaluation metric, Human error-taxonomy annotation 35 QMSum meetings, 175 annotated summaries, 4 annotators No metric correlates strongly with any error type; about a third of metric-error pairs ignore or reward the error; perplexity rewards wrong speaker references at r = +0.44
Apple 2024 arXiv technical report, 2024 (v2 2026) On-device language model, Abstractive summarisation method 2 models; ~3B on-device The 4,096 is the core pre-training sequence length, not a runtime context cap; on-device summariser built for emails, messages and notifications, 71.3 to 74.9 percent good-result ratio
Chen 2016 LREC 2016 Corpus resource, Decision and action-item extraction 22 ICSI meetings, 21,035 utterances Only 318 turns (1.5 percent) carry an actionable item; binary annotator agreement kappa 0.644, action-type agreement 1.000
Dai 2026 arXiv preprint, 2026 Far-field single-channel transcription benchmark, Joint speaker-attributed transcription metric AMI-SDM, AliMeeting, AISHELL-4 23.32 percent cpWER on AMI single distant microphone; 25.43 on AliMeeting first array channel; speaker count accuracy 76.67 percent on AMI
Golia 2023 NLPIR 2023 (ACM) Abstractive summarisation method, Decision and action-item extraction AMI corpus, human transcripts BERTScore 64.98 and ROUGE-1 36.27 on AMI gold transcripts; adding action items raises BERTScore and lowers ROUGE; action items themselves never scored on AMI
Gong 2024 arXiv preprint, 2024 Summarisation evaluation metric QMSum (232 meetings) plus 139 internal Zoom examples Language-model judges correlate with human meeting-summary scores at Pearson 0.5 on QMSum and 0.11 on real Zoom data, against 0.95 on short news summaries
Huo 2026 arXiv preprint, 2026 Far-field single-channel transcription benchmark, Joint speaker-attributed transcription metric, Speaker attribution metric AMI-SDM, AliMeeting far 24.84 percent DER and 42.55 percent cpWER on AMI-SDM; in overlapping speech every system tested was above 33 percent DER
Jones 2022 LREC 2022 Corpus resource 202 speakers, ~2,359 calls, ~840 videos Telephone and self-recorded video only; no meetings, no transcripts, no summaries; supervises speaker identity and nothing else
Kalda 2024 CHiME 2024 Workshop Far-field single-channel transcription benchmark, Array front-end signal processing NOTSOFAR-1 dev-set-2 and eval set 41.2 percent tcpWER on eval with continuous speech separation and NO diarization; distinct from the 41.4 percent published challenge baseline
Kirstein 2024b arXiv preprint, 2024 Summarisation evaluation metric, Human error-taxonomy annotation 170 QMSum Mistake summaries, 4 annotators ROUGE-LSum correlates POSITIVELY with repetition (+0.26), structure (+0.23), coreference and hallucination (+0.19) errors; best evaluator reaches Spearman -0.58
Li 2026 arXiv preprint, 2026 Far-field single-channel transcription benchmark, Near-field close-talk reference condition, Joint speaker-attributed transcription metric AMI-IHM, AMI-SDM, AliMeeting, AISHELL-4, Fisher 21.26 percent cpWER on AMI single distant microphone against 16.40 on the worn-headset mix, using an external diarization prior
F. Liu 2008 ACL 2008 short papers Summarisation evaluation metric, Extractive summarisation method 6 ICSI meetings, 36 human and 24 system summaries ROUGE-SU4 correlates with human judgement of system meeting summaries at Spearman 0.08; ROUGE-1 at -0.07
J. Liu 2023 ICASSP 2023 Decision and action-item extraction, Corpus resource AMC-A 424 Chinese meetings; AMI 101 meetings AMI's 381 action items are derived from summary links, not annotated; ICSI's 18-meeting annotation is no longer public; best English positive F1 43.12
Y. Liu 2023 arXiv preprint, 2023 Human error-taxonomy annotation, Summarisation evaluation metric 22,000 summary-level annotations, 28 systems, 3 written-text datasets The 150-hour annotation cost is for news and written chat summarisation, not meetings; reference-free human rating correlates 0.926 with input-blind preference
Pandey 2023 arXiv preprint, 2023 Contextual biasing for rare words 16k + 29k + 3,558 in-house far-field voice-assistant utterances 43.9 and 57.2 percent relative named-entity WER reduction over an un-personalised RNN-T; 37.1 and 50.0 for the prior text-only adapter
Polok 2026 arXiv preprint, 2026 Synthetic training-data generation, Joint speaker-attributed transcription metric, Speaker attribution metric 8 evaluation conditions, 2 model families Best 16.3 percent tcpWER on NOTSOFAR-1 single-channel, measured with GROUND-TRUTH diarization supplied; synthetic plus real beats real-only everywhere
Prevot 2025 SIGDIAL 2025 Corpus resource, Discourse segmentation annotation 73 French meetings, ~24 h manually annotated Human coders agree at Cohen's kappa 0.85-0.89 on discourse-unit boundaries; a fine-tuned model reaches F 0.86, plateauing at human agreement
Shapira 2025 ACL 2025 long papers Downstream-task robustness to transcription error, Abstractive summarisation method QMSum 281 instances, QAConv 2,083 questions, MRDA 1,200 utterances Summary quality drops significantly at a noise-toleration point between 0.07 and 0.3 WER, about 0.2 on average; repairing named entities is the most effective correction
Shon 2023 arXiv preprint, 2023 Downstream-task robustness to transcription error, Corpus resource 4 datasets, >20 recognition systems Word error rate and ROUGE-L correlate at Pearson -0.9 on single-speaker TED talks; the paper makes no claim about ICSI action-item sparsity
Sun 2025 Interspeech 2025 Overlapped speech detection AMI test set; AliMeeting and LibriHeavyMix for training F1 82.76 percent for overlapped speech detection on far-field AMI; overlap is 19 percent of AMI and 42.27 percent of AliMeeting
Tipaksorn 2025 arXiv preprint, 2025 Corpus resource, Far-field single-channel transcription benchmark, Near-field close-talk reference condition 114 h, 90 sessions, 86 speakers, 9 devices at 0.12-10 m Zero-shot WER 36.4 percent at 15 cm, 44.2 at 2 m, 96.3 at 3 m, 104.2 at 10 m; far-field macro 81.6 zero-shot to 49.5 fine-tuned; CC BY-SA 4.0
Vinnikov 2024 Interspeech 2024 Corpus resource, Far-field single-channel transcription benchmark, Far-field multi-channel transcription benchmark Table 1: 107 + 36 + 137 = 280 meetings Dataset paper's own table gives 280 meetings and states NO licence; dev baseline 46.8 percent single-channel vs 32.4 multi-channel vs 12.0 close-talk
Watanabe 2020 CHiME-6 challenge, 2020 Corpus resource, Far-field multi-channel transcription benchmark, Speaker attribution metric 20 dinner parties, 4 participants each Six Kinect 4-microphone linear arrays plus worn binaural mics in real homes; automatic diarization costs about 15 absolute WER points; DiPCo is a separate corpus

Note: no duplicate files. Cornell 2025 and Cornell 2024 are two different papers by an overlapping author group, the first a journal review of both challenges and the second the CHiME-8 DASR challenge description; they are indexed separately and must not be collapsed. Abramovski 2025 and Niu 2024 report the same two headline numbers (22.2 and 10.8 percent) because Niu 2024 is the system that produced them and Abramovski 2025 is the challenge summary that ranked it.

Short-cite disambiguation, because three surnames repeat in this vault. Kirstein 2024 is "What's under the hood: Investigating Automatic Metrics on Meeting Summarization" (Findings of EMNLP 2024); Kirstein 2024b is "Is my Meeting Summary Good? Estimating Quality with a Multi-LLM Evaluator" (arXiv, November 2024), by an overlapping author group. F. Liu 2008 is Feifan Liu and Yang Liu on ROUGE and human evaluation; J. Liu 2023 is Jiaqing Liu and colleagues on action-item detection; Y. Liu 2023 is Yixin Liu and colleagues on the RoSE summarisation-evaluation benchmark. The three Liu papers share no authors and no domain.

Note on disagreeing published counts for NOTSOFAR-1. The dataset paper's own Table 1 gives 107 + 36 + 137 = 280 meetings with 22 training and development speakers and 10 evaluation speakers (Vinnikov 2024). The challenge summary gives 110 + 35 + 170 = 315 meetings with 22 and 13 speakers (Abramovski 2025). A third paper's corpus-comparison table repeats 315 sessions and 35 speakers, and records the licence as CC BY-NC-ND 4.0 (Tipaksorn 2025), while the dataset paper states no licence at all. Each entry reports what its own paper says; no reconciliation is asserted here.

What far-field meeting transcription actually scores, and under which recording condition

Abramovski 2025 - the single-microphone penalty, measured on the same meetings (N=315 meetings, 30 rooms)

Cornell 2024 - the CHiME-8 DASR challenge, and why speaker counting drives the error (N=4 scenarios)

Cornell 2025 - how weakly summary quality tracks transcription quality (N=32 systems, 9 teams)

How the best-scoring systems are built, and what they require to work

Niu 2024 - the system that won both NOTSOFAR-1 tracks, and what it cost to build (N=NOTSOFAR-1 dev and eval sets)

Shi 2023 - what the extra microphones bought on AliMeeting (N=104.75 h train, 4 h eval, 10 h test)

Dai 2026 - hierarchical speaker classification bolted onto a speech language model (N=3 meeting corpora)

Huo 2026 - explicit timestamps inside a language model, and what that costs in words (N=2 meeting corpora)

Kalda 2024 - the NOTSOFAR-1 entry that skipped diarization entirely (N=NOTSOFAR-1 dev-set-2 and eval set)

Li 2026 - handing the language model a diarization prior instead of asking it to diarize (N=5 test sets)

Pandey 2023 - pronunciation-aware biasing for personal names, measured on a voice assistant (N=3 test sets)

Polok 2026 - what simulated conversations buy, and the oracle-diarization caveat under the numbers (N=2 tasks)

Sun 2025 - detecting when two people talk at once, on far-field meeting audio (N=3 corpora, AMI test set)

What meeting summarisation is, and what it has been built and scored on

Kumar 2022 - a survey and leaderboard of meeting summarisation systems (N=over 40 papers surveyed)

Rennard 2023 - the survey, and the corpus scarcity behind every number in it (N=3 corpora, about 280 h)

Golia 2023 - action items folded into the summary itself, scored on gold AMI transcripts (N=AMI corpus)

J. Liu 2023 - the action-item corpus that exists, and the two that do not (N=424 Chinese meetings, 101 AMI)

Shapira 2025 - how much transcription error a downstream task survives, and which errors matter (N=3 tasks)

Shon 2023 - a spoken-language benchmark, and what it does and does not say about meeting corpora (N=4 tasks)

Whether a summary or an action item can be scored at all

Kirstein 2024 - human annotators against nine automatic metrics (N=35 meetings, 175 annotated summaries)

Gong 2024 - language-model judges fail on meeting summaries, measured against human scores (N=2 datasets)

Kirstein 2024b - a multi-agent evaluator scored against human error annotation (N=170 meeting summaries)

F. Liu 2008 - ROUGE against human judgement on meeting summaries, and how to make it less bad (N=60 summaries)

Y. Liu 2023 - what a rigorous human summarisation protocol costs, and what it was applied to (N=22,000 annotations)

What the corpora actually contain, and which of the three outputs they can supervise

Chen 2016 - ten assistant actions annotated onto 22 ICSI meetings, and how rare they are (N=21,035 utterances)

Jones 2022 - a telephone and video corpus for speaker recognition, with no meetings in it (N=202 speakers)

Prevot 2025 - a French meeting corpus segmented into discourse units, and what that costs (N=73 meetings)

Tipaksorn 2025 - the same conversation on nine devices at measured distances, 0.12 m to 10 m (N=114 hours)

Vinnikov 2024 - the NOTSOFAR-1 datasets as the dataset paper itself describes them (N=280 meetings)

Watanabe 2020 - twenty real dinner parties on six Kinect arrays, and what diarization costs (N=20 parties)

What a phone can actually run, and what it was asked to summarise

Apple 2024 - the on-device foundation model, its 4,096, and the summaries it was built for (N=2 models)

Summary Notes for Platform Work

The single-microphone penalty is measured, and it is roughly a factor of two. On the same 170 real office meetings, the same team's system scored 22.2 percent time-constrained speaker-attributed word error rate from one distant microphone stream and 10.8 percent from a seven-microphone tabletop array (Abramovski 2025, Niu 2024). The published baseline for the same single-channel condition was 41.4 percent. The mechanism behind the gap is named in both papers and is not fixable with a better model on one device: spatial information across microphones is what lets a system separate two people talking at once, and one device does not have it. This number is the vault's most direct evidence bearing on the charter's phone-is-the-microphone constraint, and it is also the number a competitor selling a hardware recorder would quote back. Note that all these measurements come from the far-field array device family and none from a handset, so the phone's actual position on this scale is unmeasured, not favourable.

Half the error a user would see in an attributed transcript is attribution, not misheard words. On the same far-field meeting audio, speaker-agnostic character error was 17.3 and 18.4 percent while speaker-attributed character error was around 30 percent (Shi 2023, Far-field multi-channel transcription benchmark), and the NOTSOFAR-1 winner's speaker-agnostic score was 17.7 percent single-channel against a speaker-attributed 22.2 percent (Abramovski 2025, same family). For a product whose promise is a summary and action items with names attached, the attribution half is the half that shows.

Improving the intermediate metric can make the product worse, and this is reported three separate times. A diarization stage that improved diarization error rate from 23.43 to 21.07 percent made transcription worse, from 12.87 to 14.13 percent (Niu 2024). Joint fine-tuning cut speaker-attributed character error by 34 to 37 percent relative while the separated audio's signal-to-noise ratio fell by roughly 19 dB (Shi 2023). Across 22 challenge systems, diarization error rate predicts transcription error well overall but the correlation collapses to about 0.16 to 0.21 among the top systems (Cornell 2025). The practical reading for this project is that a dashboard of component metrics will mislead, and only an end-to-end measurement of the three delivered outputs is worth trusting.

A summary appears to survive a transcript that would look bad quoted as a number, but the evidence is automatic scores, not readers. On the NOTSOFAR-1 scenario, systems above 50 percent tcpWER produced summaries scoring roughly on par with a system near 11 percent, and the correlation between transcription error and the best summary metric was only PCC -0.51 (Cornell 2025). This is the strongest positive signal in the vault for the charter's thesis, because it means the deliverable the customer sees may be more robust than the raw recognition quality suggests. It is also the vault's most citable-in-the-wrong-direction finding, because every one of those summary scores came from a language model judging another language model's output, no human read any summary, and the authors themselves say the metrics are unreliable. This bullet crosses two mechanism families deliberately, Joint speaker-attributed transcription metric and Summarisation evaluation metric, and the crossing is the finding: they do not move together.

Nothing published measures a summary written from recognised speech, except the one experiment above. Every meeting-summarisation benchmark number in the vault, up to 56.26 ROUGE-1 on AMI and 60.7 on ICSI, was computed on transcripts that were human-produced or human-corrected precisely so that recognition errors would not compound (Rennard 2023, correcting a contrary claim in Kumar 2022). So the field's headline summarisation scores and this project's actual product are two different measurements, and the project must never cite the first as evidence for the second. Only three English meeting corpora with summaries exist at all, together about 280 hours.

The metrics that score a summary do not separate a good one from a fluent wrong one. Across nine metrics against expert human error annotations, none correlated strongly with any error type, about a third of the metric-error pairs ignored or rewarded the error, perplexity rewarded wrong speaker attribution at +0.44, and BLEU and QuestEval rewarded hallucination severity at +0.35 and +0.34 (Kirstein 2024). Independently, a peer-reviewed survey cites nearly 30 percent of neural sequence-to-sequence summaries suffering fact fabrication (Rennard 2023). Any quality claim this project makes about its summary or its action items will have to rest on a rubric and human raters; there is no number available to buy it with.

Action-item extraction is close to unmeasured. It is a named sub-task with its own dialogue-act-based literature and its own separately annotated section in the AMI and ICSI reference summaries (Rennard 2023), but no paper in this vault reports a benchmark score for it, no error taxonomy in Kirstein 2024 scores it, and the one downstream summarisation experiment folded action items into a single 200-word summary prompt and never scored them separately (Cornell 2025). The charter names action items as one of three deliverables; the literature offers no way to say how good they are.

Nothing in this vault demonstrates minutes-after-the-meeting delivery, and the one figure that exists points the wrong way. The most efficiency-oriented system across both CHiME challenges, built specifically to be practical and using no ensembling and no diarization refinement, still ran at a real-time factor above 2 (Cornell 2025). The NOTSOFAR-1 winner reports about six hours of GPU time to decode one development set (Niu 2024). Neither is a product engineering effort and both are unoptimised research pipelines, so this is not evidence that the charter's finished-before-the-walk-back constraint is unreachable; it is evidence that published accuracy and published speed have never been demonstrated together, and that the project cannot cite a research word error rate and a product latency in the same breath.

Real matched training data beats architecture, and it is the lever a small team actually has. The largest single improvement reported anywhere in the vault came from fine-tuning the recognition model on real meeting audio processed the same way the deployment audio would be: 16.57 to 9.87 percent tcpWER, a 40 percent relative cut, against 8.46 to 7.50 percent from every architecture modification the same team made (Niu 2024, confirmed in Abramovski 2025). The same effect appears in diarization, 21.51 to 16.52 percent DER. For a project whose capture condition is a specific device in a specific pose, this says the defensible asset is condition-matched data, not a better model.

The 2026 frontier for one distant microphone is roughly one word in five, and it takes two models to get there. Three independent 2026 systems evaluated on the AMI Single Distant Microphone condition, one tabletop microphone in a meeting room, report attributed word error rates of 42.55 percent (Huo 2026), 23.32 percent (Dai 2026) and 21.26 percent (Li 2026). The best of them reaches that figure only by taking speaker labels and segment boundaries from a separate trained diarization system and giving them to a language model as a prior; the diarization stage alone contributes 13.72 percent diarization error. Measured inside that same system, moving from worn headset microphones to the single tabletop microphone costs 4.9 absolute points of attributed word error, 16.40 to 21.26 (Li 2026). None of the three reports latency, memory or real-time factor, so the charter's finished-before-the-walk-back constraint is untested against any of them. A fourth 2026 paper reports 16.3 percent on NOTSOFAR-1 single-channel (Polok 2026), which looks like a further improvement and is not: that number is computed with ground-truth diarization supplied, and the paper says so in one sentence that appears in no table caption.

Distance, measured directly, is worse than the meeting-corpus numbers suggest. The only study in the vault that puts several single-channel devices at measured distances on the same conversation reports word error rates of 36.4 percent on a table condenser at 12 to 15 cm, 44.2 percent on a tabletop unit at 2 m, 96.3 percent at 3 m and 104.2 percent at 10 m from an off-the-shelf model, falling to 20.4, 26.4, 58.2 and 64.0 after in-domain fine-tuning (Tipaksorn 2025). The language is Thai, so the levels do not transfer to English, but the shape does, and it is corroborated in English by the NOTSOFAR-1 baseline's 12.0 percent from close-talk microphones against 46.8 percent from a single device stream on the same 36 meetings (Vinnikov 2024). Two engineering results from the same study transfer without qualification because they are about acoustics rather than language: classical dereverberation and denoising front ends made recognition WORSE at every distance, and a model fine-tuned on one microphone type collapsed on every other type unless reverberation augmentation was added. For a product whose capture device and pose vary across users, the second of those is the binding constraint.

"Single-channel" in this literature is not what a phone application gets. The NOTSOFAR-1 dataset paper states plainly that its single-channel condition is a commercial conference-room device's output stream after that device's own echo cancellation, dereverberation, beamforming and noise suppression, and it flags the distinction from single-microphone audio as important enough to give its own subsection (Vinnikov 2024). Every NOTSOFAR-1 single-channel figure this project might quote, including the 22.2 percent winning score (Abramovski 2025), was therefore measured downstream of hardware signal processing. The only vault entry whose single-channel condition is one literal microphone element is LOTUSDIS (Tipaksorn 2025), and its far-field numbers are far higher. This distinction moves in the wrong direction for the charter's phone-is-the-microphone bet and it is invisible in every number as normally quoted.

A summary survives noise better than a transcript deserves, but only two of four models survived 25 percent, and the audio was synthetic. The threshold result the project has been carrying reads, in the authors' own words, "models tend to tolerate a noise level of about 0.2 WER (NTP between 0.07 and 0.3)" (Shapira 2025), with per-model points of 0.070, 0.195, 0.249 and 0.297. The conditional half of the claim is the better-supported half: repairing only the named entities in a damaged transcript was the single most cost-effective correction for summarisation, and a transcript at 0.9 word error rate repaired down to 0.4 by fixing content words produced summaries much preferred over a differently-damaged transcript at the same 0.4. Two things bound it. The audio was text-to-speech synthesis reverberated and noised in software, with no overlapping speech and no real microphone, and the summaries were ranked by a language model rather than read by people. The one measurement in the vault that used real meeting audio found a much weaker coupling, PCC -0.51 (Cornell 2025), while the one that used single-speaker TED talks found a much stronger one, Pearson -0.9 (Shon 2023). The three do not agree, and the axis they differ on is the audio condition.

Action items are the least measurable of the three deliverables, and the resources say why. There is no public English meeting corpus with human-annotated action items: AMI's widely cited 101 meetings and 381 action items are DERIVED by treating dialogue acts linked to the action-related section of the abstractive summary as positive, and ICSI's 18-meeting annotation is no longer publicly available (J. Liu 2023). The largest purpose-built corpus is Chinese, 424 meetings, and its inter-annotator agreement is Cohen's kappa 0.47. The best reported English positive F1 is 43.12, on human transcripts. Independently, the one ICSI annotation that is public covers 22 meetings and finds that only 318 of 21,035 turns carry an actionable item, 1.5 percent, with binary annotator agreement of kappa 0.644 (Chen 2016). Set that against kappa 0.85 to 0.89 for annotating discourse-unit boundaries on the same kind of speech (Prevot 2025): structure is decidable and importance is not, and the gap is a property of the task, not of the annotators.

Nothing has changed since 2008 about whether a meeting summary can be scored automatically, and four papers now say so. ROUGE's rank correlation with human judgement of machine-generated meeting summaries was 0.08 in 2008 (F. Liu 2008). In 2024 ROUGE-LSum correlated POSITIVELY with the presence of repetition, structure, coreference and hallucination errors on meeting summaries, meaning it rewards them, and BERTScore ran between -0.32 and +0.08 (Kirstein 2024b). Language-model judges, the modern replacement, correlate with human scores at Pearson 0.5 on QMSum and 0.11 on real Zoom meeting data, and assign perfect completeness scores to machine summaries shown to them as their own reference (Gong 2024). On written news and chat summarisation the same metrics behave far better, reaching system-level Kendall 0.84 (Y. Liu 2023), which is why the failure is specific to meetings and why borrowing a written-text evaluation result would hide it. Note also that Y. Liu 2023's "over 150 hours for 22,000 annotations" is the cost of annotating news articles and written chat logs, not meetings.

The platform vendor's on-device model has never been asked to do this job. The only evidence in the vault about what a phone can run is a roughly 3-billion-parameter model quantised to about 3.7 bits per weight, whose summarisation feature was specified, trained and evaluated for emails, messages and notifications (Apple 2024). It needed a fine-tuned adapter before it would follow a summary specification at all, its adapter training summaries were synthetic output from a larger server model, and the report names prompt injection as a live failure: the model follows instructions found inside the content it is meant to summarise, which a meeting transcript will contain by accident. The 4,096 figure this project has been carrying is the core pre-training sequence length and, separately, the pre-training batch size in sequences; the same report describes later training stages at 8,192 and 32,768 tokens, and states no runtime context cap anywhere. No latency, throughput or memory figure appears in the report.

View as PDF Download PDF