CopunditSample
View as PDF Download PDF

Brainstorm, the adjacent ideas

The ideas a panel of experts generated for this project, each one scored and each one audited against the evidence this project has actually gathered. Seven specialists were given the same question and the same approved evidence base, proposed independently without seeing each other, and then scored every surviving idea on four measures. A separate reader, who proposed nothing of their own, checked each idea against the evidence and recorded exactly where it rests on something nobody has measured. Ideas that contradict this project's own rules were kept and argued rather than smoothed away, and every idea that was folded into another one is listed at the end, so nothing generated here is hidden.


Contents

The question put to the panel

The evidence floor is approved and the pitch has been re-scored against it. Two of its five claims are dead: the phone's microphones are not reachable by an application as an array, and the person buying the 159 dollar recorder gives reasons an application cannot supply. What adjacent ideas does this evidence open? Propose ideas that stay inside this project's scope constraints where you can, and argue a pivot openly where the evidence points outside them. For each idea: what it is in one paragraph, which use-case bucket it sits in, who pays for it, and what would kill it.

How to read the scores

Novelty is how far the idea sits from what anyone in this market is already selling. Defensibility is how hard it would be for a competitor, or for the operating system underneath the product, to copy the idea or simply absorb it. Time-to-market is how soon it could be in a user's hands given what the evidence says is already possible today. Strategic fit is how closely it serves the customer and the promise this project was started on. Every total is the sum of four averages taken over seven independent specialists who scored the same slate without seeing each other's marks; the averages were counted mechanically, and the order inside each table is by total and has not been adjusted.

What the audit found before anything was scored

Nothing on this slate came back green. Green would mean an idea whose mechanism, whose payer and whose kill test all rest on evidence that has already been established, and the auditor's finding is that such an idea is not currently constructible in this project: the data inventory returns nothing on all five of the claims the pitch needs data for, and every company figure in this field carries a footnote while every customer figure is borrowed from a neighbouring market. Forty of the forty-one ideas came back yellow, which means the idea stands but rests on at least one thing nobody has measured, and in most cases the seat that proposed the idea named that gap first, which the auditor calls the round's strongest quality signal. Red means the idea contradicts something this project has already investigated and found false.

One idea was red. A founder-seat proposal to have the phone speak the recording notice out loud claimed that the two-tap interaction survived intact because the announcement is fired by the start tap and is therefore not a third one. The competitive map and the evidence base both state, in the same words, that the compliant interaction is at minimum start, announce, stop, and the consent-law work refuses that exact re-labelling in advance: the announcement is a spoken act in a room full of colleagues, so the cost is the utterance, not the finger on the glass. The idea relocated the cost to the machine and then declared it gone, which makes it an undeclared contradiction of this project's own two-taps rule. The mechanism itself is not dead, and three other seats propose it while pricing the cost honestly.

Underneath the slate the auditor found two shared assumptions rather than forty-one independent ideas. Six ideas from six different seats, with a seventh losing half its argument alongside them, all rest on one unverified proposition: that a recording fails silently often enough, on the handsets this customer actually holds, to be worth buying insurance against, and that the failure is detectable from inside a running application. Both circulating failure rates for that have been struck from this project's record as unsourced, no shipped recorder has published a defect rate against which any target could be judged, and the specific case where the microphone keeps delivering buffers full of digital silence may not be detectable from inside the application at all. The second shared assumption sits under a different seven ideas: each takes a percentage from a sample of sixty-four stated purchase reasons as a statement about demand, when that sample was assembled from forums of people who already own the hardware, so it measures why buyers stayed with a device rather than what a buyer still choosing would do. Every percentage on the slate inherits that selection. Read together, the slate is closer to two bets than to forty-one ideas, and both bets are one measurement away from resolution in either direction.

The same product, better

ID Novelty Defensibility Time-to-market Strategic fit Total
S2 (folds M1) 3.9 2.3 4.4 4.9 15.5
P5 4.0 3.0 3.7 4.7 15.4
S3 4.1 2.0 4.9 3.9 14.9
M3 4.6 3.0 2.9 4.3 14.8
M6 (folds T3) 3.9 2.4 3.0 5.0 14.3
S1 5.0 2.3 2.7 4.0 14.0
V1 (folds P3) 4.0 2.7 4.1 3.1 13.9
V2 2.7 3.0 4.4 3.7 13.8
P4 4.3 1.7 3.0 4.7 13.7
G2 4.0 2.1 3.0 4.4 13.5
T5 2.6 1.1 5.0 4.7 13.4
T1 3.3 1.1 4.7 3.3 12.4
T2 2.3 1.0 5.0 3.6 11.9

S2 (folds M1) , The recording-is-alive watchdog

The evidence base names the silently lost recording session as the only mechanism in this round that is both buildable and unclaimed, and no capture-defect rate has ever been published for any shipping recorder on either platform. The self-check runs on three signals the speech field already uses: runs of exactly-zero audio samples, which a live microphone carrying its own dither and room noise floor cannot produce; a long-window noise-floor and spectral-flatness track, because a microphone that is live but muffled in a bag or face down loses high-frequency energy and gains cloth-rustle transients; and a speech-presence rate from an always-on voice activity detector of the Silero class, about 2 MB in size. What reaches the user is a warning rather than another tap, so the two-tap interaction holds, and the detector is specified as false alarms per recorded hour at a fixed miss rate on true failures rather than as an accuracy, because the quiet room that is recording perfectly well dominates deployment time. The technology survey already fixes the study shape, at least 500 sessions from at least 50 real users across both platforms with a pass above 99.0 percent defect-free completion, and notes there is no category baseline to judge that bar against; folded into this entry is a pre-flight probe that refuses to accept a meeting until it has proved, on this handset and this operating-system build, that a session survives a screen lock, an application switch and a simulated interruption, and that reports the exact setting that fixes it rather than reprinting the community sites' multi-step instructions.

P5 , The recording that tells you it is still recording

The competitive map rates trust death by a single lost recording as this project's highest risk of any pattern in the document, and the technology survey supplies the failure modes: on iOS the privacy subsystem refuses a backgrounded application's attempt to restart the microphone after a call, returning OSStatus 561145187; developers report an input tap that keeps firing while delivering zero-filled buffers, so the application writes a valid-looking file of digital silence; and on Android four named handset manufacturers kill the recording service regardless of correct interface use. The evidence base records that nobody has measured capture survival on any handset, that both failure rates circulating in the field are unsourced, and that no vendor publishes a guaranteed maximum time to a finished summary, so the two things a buyer of insurance wants, does it survive and when does it finish, are both unclaimed. The product continuously verifies that audio is really arriving, escalates to the user the moment it is not, writes through to storage in short increments so a hard kill costs seconds rather than the session, and publishes both of the numbers it measures. A platform owner will never occupy this position, because it would have to ship a feature whose whole content is a warning that its own operating system may take the microphone away, and it will never publish a survival rate for its own audio stack.

S3 , Publish the entity error rate, then move it

No named-entity word error rate on far-field single-channel meeting audio exists in any language on any corpus, and everything the pitch leans on downstream is conditional on that missing number. The finding that summarisation tolerates transcription errors, where systems above 50 percent time-constrained speaker-attributed word error rate produced summaries scoring roughly on par with systems near 11 percent (arXiv 2507.18161), holds only when the errors fall on fillers, false starts and conjunctions, while far-field room acoustics smear consonants so that proper nouns, numbers and acronyms fail first, and those are precisely the words an action item is made of: who, by when, which account. Reporting the number costs no new audio, because the AMI and AliMeeting meeting corpora both ship word-level verbatim transcripts alongside a close-microphone reference, so entity spans can be marked on the reference and scored against the far-field decode. The product lever that moves the number is already shipping rather than research, contextual biasing on a user's personal entities in the PROCTER line of work at a reported 44 percent relative improvement overall and 57 percent on rare personalised entities, carried here as a direction and never as a number because it states no audio condition and no test set, and the cost is that a privacy-positioned product starts reading contacts and calendar.

M3 , Seat-position speaker labels from the ambisonic channel

The technology survey records that iOS 26 adds capture of First Order Ambisonics (FOA), a four-component recording that encodes the direction each sound arrived from, through the standard media-writing interface, and read carefully this does not contradict the platform finding that killed the microphone-array claim, because it is not raw per-microphone audio: the phone computes the encoding itself and what it releases is a spatial description rather than the individual capsule signals. In a meeting people keep their seats, so a direction is a stable per-speaker key for the length of a session, and clustering utterances by direction is a route to the one capability the platforms withhold, since Google's own recorder has labelled speakers since its version 4.2 and neither platform exposes speaker separation to anyone else. Two engineering conditions travel with it: the audio session must not declare a voice-communication mode, because the system then inserts voice-processing units whose automatic gain control and noise suppression destroy the spatial information, and the path is iPhone-only on iOS 26, so the Android product gets speaker labels some other way or does without. The technology survey calls this measurement the highest-value trigger in the document and states that nobody has made it in either direction.

M6 (folds T3) , A deadline that can be published

The technology survey records that the latency intuition is inverted: data-centre silicon is roughly forty to sixty times faster at raw inference than a recent phone, and yet cloud batch products measure slower end to end, because queueing and transport rather than computation are what the user waits for. If recognition is finished at the moment the user taps stop, the wall-clock wait stops being a function of a queue and becomes a function of transcript length, which is bounded, measurable and therefore publishable, and the evidence base fixes what is unclaimed here: several vendors publish typical processing times and none publishes a guaranteed maximum. The cost is carried with the idea rather than around it, since the phone's on-device language model holds 4,096 tokens of context against an hour-long transcript of eleven to thirteen thousand, which makes the post-tap pass a chunk-and-merge whose quality nobody has evaluated. Folded in from the go-to-market seat is the move that turns an engineering property into a reason to switch at the moment of purchase, putting the number into the terms of sale with a remedy behind it rather than into the marketing copy, sold away from the app store because inside it the store is the merchant of record; and the capture-integrity dividend only this seat noticed is that a session that looks healthy while producing no words for a sustained window is either a silent room or a revoked microphone, and the product can ask which.

S1 , Ambisonic front end on the iOS 26 capture path

First Order Ambisonics (FOA) is a four-component description of a sound field, one omnidirectional pressure channel plus three pressure gradients along three perpendicular axes, and iOS 26 added FOA capture through the standard media-writing interface, which is the one route to a spatial front end that does not need the raw per-microphone access both platforms withhold. It matters because guided source separation, the 2018 statistical front end that every top system in two consecutive CHiME challenge rounds still uses, fits a spatial covariance matrix from the phase difference between microphones and is mathematically undefined on a single channel, whereas from the pressure and gradient channels an application can compute a direction-of-arrival vector for every time-frequency bin, and the parametric spatial audio line of work builds beams and separation masks from that four-component format alone. The size of the prize is the one measurement the field has made on identical audio, 22.2 percent single-channel time-constrained speaker-attributed word error rate against 10.8 percent with a microphone array on the 170 blind NOTSOFAR-1 evaluation meetings (arXiv 2501.17304), which is one system's score on one evaluation set and not a bound on anything. Two conditions travel with it, never activating a voice-communication audio session because the system then inserts voice-processing units whose automatic gain control and noise suppression destroy the phase the direction estimate reads, and verifying the sample rate actually delivered, and the honest exposure is that the path is iPhone-only and gated on iOS 26.

V1 (folds P3) , The announcement inside the recording

On the start tap the application opens capture and shows a one-line script for the user to say out loud, then listens for that script inside its own captured audio and pins the detected utterance as the first anchored segment of the transcript, writing the time, the detected wording and the recording identifier into a consent record bound to the file. The consent-law research reports Washington as the one jurisdiction where a captured announcement is the compliance mechanism itself rather than proof of it; everywhere else it is the evidence no defendant in this category holds, and its absence is the whole complaint, since the competitive map states that a two-tap product with no bot in the room and no announcement is exactly the Chamberlain v. Granola fact pattern, brought against a product that was already processing on the user's own device. The binding in-person all-party set is twelve jurisdictions, thirteen on the cautious Missouri reading, with Nevada out and Oregon and Hawaii in, and the strictest-standard rule attributed to Kearney v. Salomon Smith Barney makes all-party consent the national operating default, so the move is to stop treating the announcement as a legal chore bolted onto the interaction and sell it as the deliverable, the only recorder that hands you a record showing you told the room. Folded in is the absorption argument: a platform owner structurally cannot ship this, because ambient recording that announces itself, logs who was told and keeps that record next to the audio puts the shipper inside the wiretap liability chain for its entire install base, which is what Basich v. Microsoft is, so Apple and Google ship a recording indicator and stop there.

V2 , Nothing outlives the summary

When the three outputs are produced the audio is destroyed, and the speaker embeddings, the numerical voice vectors the speaker-separation step produced in order to tell one talker from another, are destroyed with it, so no vector ever crosses a recording boundary and no enrolment store exists anywhere in the system. The technology survey draws the legal line in exactly this place: anonymous ephemeral clustering inside a single recording is the described safe harbour, while tying a voice vector to an identity by enrolment, a calendar invitation or an email address is the transition into named biometric identification, which is the specific allegation in Cruz v. Fireflies.AI Corp. and Basich v. Microsoft Corp. under the Illinois Biometric Information Privacy Act. The commercial precedent is compliance as distribution: Jamie holds ISO 27001 and SOC 2 Type II certification, keeps European-only servers, deletes audio the moment the transcript exists, contractually bars model training, and sells at two to three times the category price on that posture. Nothing is given up by deleting, since selling or licensing the recorded audio is recorded as a structurally closed way to make money here, and the General Data Protection Regulation (GDPR) Article 17 duty to isolate and permanently delete one named individual's voice and contributions on demand is trivially satisfied by a system that kept neither.

P4 , Record before install, pay off-store

The survey of the existing application shelf counts a seven-step first-run gauntlet for any third-party application, launch, account creation, biometric prompt, microphone permission, notification permission, dismissing the subscription upsell, then record, against one pinch on the dedicated device, and finds that returning users already reach one to two taps on Granola through a watch complication, a Live Activity, the Dynamic Island or an Android Quick Settings tile, and on Pixel Recorder through a widget, concluding that two taps is table stakes for a returning user, unreachable for a new one, and that the first-run path is the one genuinely unclaimed position on the shelf. Every surface that compresses that path is owned by the platform and handed to developers at no charge, and this category uses those surfaces only for people who have already installed. The shape is that the first recording starts from a link, a code or a system action with no install and no account, and the account is asked for at the moment the three outputs are handed over, which is also the moment the buyer has seen the product work. That inverts the crowding pattern the competitive map describes, where store ranking and advertising decide and acquisition cost climbs past what a subscriber is worth, and it puts the purchase on the seller's own checkout so that one store review cannot close every route at once.

G2 , Unmetered capture, priced per finished meeting

Every free tier in this field meters the thing the category is bought for and the thing that is about to cost nothing: transcription dominates marginal cost at 0.21 to 0.62 US dollars per recorded hour while the summarisation call sits at 0.002 to 0.10, two to four orders of magnitude below it, and moving recognition onto the phone takes the metered half to roughly 0.05 an hour or to zero, which Apple's on-device speech framework from iOS 26 and Android's on-device generative stack make available to a third party now. So the packaging the whole shelf shares is inverted: capture and the transcript are never metered, never capped and never lost behind a paywall, and money attaches to the finished summary and action items delivered under the deadline, which is the moment the value lands. That fixes the loss-making shape the funding map works through, where 100 free signups consume up to 200 US dollars of inference a month while the three to five who convert generate 45 to 75, and where the field's other three answers each churn exactly the users you want: minute caps (Otter at 300 a month, Wave at 30, Notta at 120 with a three-minute file cap), history caps (Granola at 30 days) and summary caps (Fathom's first five calls). The counter-evidence is not smoothed away: zero marginal cost is not a moat, it is the mechanism by which the price of the core function goes to nothing, which is the argument for putting the price on the delivery rather than on the recording.

T5 , The remorse window

This is the one segment whose willingness to spend on this job is demonstrated rather than borrowed and it is publicly reachable, because the funding map records retail and marketplace listings as where the category's public buyer evidence lives, in Amazon and Best Buy verified-purchaser reviews, and the hardware vendors run a 30-day no-questions return policy, which means they model remorse and will not disclose the rate. The channel is timed by intent rather than sorted by demographics: acquire on the return, replacement and alternative queries, and on comparison content aimed at somebody holding a device they have decided against, at the one moment when 159 dollars is back in their pocket and the problem that made them spend it is still not solved. The same review corpus supplies the vocabulary to write with, since the sixty-four stated purchase reasons this project works from were pulled out of it. The discipline that makes the channel affordable rather than a money pit is that 28.1 percent of those statements are about capturing a cellular telephone call, which no application on either platform is permitted to do, so the copy has to disqualify that buyer at the click rather than at the refund.

T1 , Sell readiness, meter the hours

The category prices flat at 8 to 19 US dollars a month for an individual tier against a cost metered by the recorded hour at 0.21 to 0.62 US dollars assembled, dominated by transcription with the summarisation call two to four orders of magnitude below it, and on the funding map's own worked example a 15 dollar subscription netting about 13 after commission breaks even near 32.5 recorded hours a month, so a user recording two to three hours a working day costs the vendor more than they pay. That inverts the standard consumer playbook, since the engaged user is the loss and the dormant subscriber is the margin, and every growth tactic that drives usage makes the accounts worse. The packaging that matches the cost charges a standing fee for readiness, meaning the recorder is armed, the guarantee is live and the archive is kept, and meters the recorded hour on top of that in packs that do not expire, so the heaviest user becomes the best customer instead of the worst one. The evidence that it could work is that the whole shelf already meters privately and incoherently, with seven vendors each capping a different quantity, 300 minutes a month, 100 a week, 120 with a three-minute file cap, 30 a month, 800 lifetime, five summaries, 30 days of history, and nobody has put the meter on the price list where the buyer can see it and buy more of it, while the hardware buyer already lives under one, since the 159 dollar device buys 300 transcription minutes a month and no more.

T2 , One receipt a year, issued by us

The funding map names the hardware rail's real advantage in one line, that the cash arrives before the service does, which is how a company that has shipped over two million units has raised under 6 million US dollars of disclosed funding, and the open question underneath it is what pays for a software entrant's acquisition once the 159 dollars is removed. The only answer available to software is the same one the device uses, which is to take the year in a single transaction at the moment the buyer was already reaching for a card. The second half is the receipt itself: the same map describes the device buyer as an individual who frequently expenses the purchase without going through their organisation's technology function, and records Japan's tax code instantly expensing a recorder priced under 100,000 yen, a demand subsidy that reaches objects and not software, so an annual invoice issued by our own merchant of record and carrying the employer's name and tax registration is the artifact that clears the same expense system a device receipt clears. One correction already on this project's record governs how any of this may be said out loud: annual prepay is a cash-timing and receipt mechanism, never a discount claim, and an annual figure is never set against anybody's month-to-month figure.

An adjacent product for the same buyer

ID Novelty Defensibility Time-to-market Strategic fit Total
S5 (folds V3) 3.7 2.3 4.1 4.0 14.1
G4 (folds T4) 3.4 2.7 2.7 3.3 12.1

S5 (folds V3) , Action items that abstain, with the audio behind them

Action-item extraction is the one output in this category with no accuracy number in any product or any paper, and its named failure mode is a fluent, plausible task attributed to somebody who never agreed to it. The failure modes are asymmetric, since a missed commitment leaves the user where they started while an invented one actively misleads them, so precision is weighted far above recall and the product surface that expresses that operating point is abstention: below a confidence threshold the system reports an unassigned commitment rather than naming a person, and every item it does assign carries a timestamped span the user can play. Assigning owners from what was actually said in the room, and abstaining when unsure, avoids voiceprint enrolment entirely, which the technology survey describes as the safe harbour against named biometric identification under the Illinois Biometric Information Privacy Act at 740 ILCS 14/15(b), and folded in is the counsel's framing that a wrong name is inaccurate personal data about a person who never installed anything, carrying an accuracy duty under GDPR Article 5(1)(d) and a rectification duty under Article 16 owed to the person recorded rather than to the user who recorded them. The supervision exists, which is unusual on this slate: the AMC-A corpus carries 1,506 sentence-level action items over 424 meetings at 0.47 annotator agreement, AIMU 318 actionable turns over 21,035 turns of 22 ICSI meetings, AMI 381 items over 101 meetings, and the Kirstein set 175 machine summaries with human error-type labels at Krippendorff's alpha 0.76 to 0.83, the only material annotated for the exact error this suppresses, while the automatic metrics reward that error, with perplexity at +0.44 on wrong speaker references and BLEU at +0.35 on hallucination (arXiv 2404.11124).

G4 (folds T4) , The recap the room receives

The bot competitors get the attendee list free from the calendar invitation, which is exactly why Otter and Fireflies can send a recap after a call and an in-person recorder cannot, and the meeting this project is about has no invitation, no attendee list and no shared artifact, which is the structural reason its notes die inside one person's application. So the outgoing recap is addressed to the people who were in the room, assembled from action items that carry owners, and identifying the room becomes the product's real input rather than an afterthought. The supervision for the hard half is thin enough to be worth naming: the AMC-A corpus carries 1,506 sentence-level action items over 424 meetings and 306,846 utterances at 0.47 annotator agreement and the AMI corpus supplies 381 items over 101 meetings, while the technology survey records that no benchmark anywhere scores action items together with their owners, so this output is unmeasured across the whole field and an entrant is not behind on it. Folded in from the go-to-market seat is the observation that the legal cost and the growth loop are the same action, since the announcement is paid for anyway under a binding all-party set of twelve jurisdictions and possibly thirteen, and each delivered recap is a free impression on somebody who attends the exact kind of meeting the product is for, against a fallback acquisition cost the funding map can only quote as borrowed at 20 to 40 US dollars with no footnote behind it; the honest hazard is that sending somebody a recap is proof you recorded them.

A different buyer for the same mechanism

ID Novelty Defensibility Time-to-market Strategic fit Total
M2 3.9 3.4 4.0 1.9 13.2
V6 3.1 2.3 4.0 2.7 12.1
S4 + M5 5.0 4.0 1.0 1.9 11.9
P1 3.6 4.0 1.7 1.9 11.2
V4 3.1 2.9 2.4 2.0 10.4
G5 3.0 3.0 2.0 2.0 10.0

M2 , Crash reporting for the microphone

Every product whose value depends on a long microphone session has this problem and none of them can see it, so the layer drops into an application and owns the parts that get people paged: declaring the background audio mode on iOS and the microphone foreground-service type on Android 14 and later, holding the session, subscribing to interruption, route-change and thermal notifications, restarting the service after a memory kill, requesting the battery-optimisation exemption, forcing a universally supported capture rate so that fragmented hardware does not fail silently, and emitting one structured record per session with the audio itself never leaving the device. What the buyer gets that does not exist today is a defect-free completion rate for their own installed base, broken down by handset model and operating-system build, so that a support ticket saying "it stopped recording" becomes a row in a table. The technology survey records that an exhaustive search of developer post-mortems, engineering blogs, mobile-systems literature and public bug trackers found no published capture-survival measurement on either platform. The instrument creates the category baseline as a side effect of being sold, and a public per-handset survival table is the closest thing this project has to the self-liquidating acquisition asset that the 159 dollar object gives the hardware incumbent.

V6 , The room where the announcement is already free

Same two-tap capture through the phone's own microphones, aimed at the buyer whose job already requires them to say "I am recording this" out loud before anything else happens: the recruiter running a screening conversation in a room, the reporter taking an on-the-record interview, the adjuster taking a statement, the field researcher taking informed consent. For every other buyer in this category the announcement is a third interaction and a social cost, which is the objection that makes the compliant shape expensive; for this buyer it is a step they already perform, so the cost is zero, and the artifact it produces, a timestamped in-recording consent stamp bound to the file, is something they currently keep by hand or not at all. The buyer evidence is not overstated: the sixty-four-statement sample was assembled from hardware-owner forums including the sales and consulting communities, so these roles are visible in it, but the stated reasons run the other way, with 17.1 percent naming social discretion and bot avoidance, which is the opposite posture to announcing. What does support the segment is that after the strictest-standard rule every recorder in this market faces the announcement problem, and this is the one buyer for whom it is already solved before the product arrives.

S4 + M5 , The phone-condition recording programme and the rig that makes the corpus possible

The data inventory returns nothing on every claim that needs data, and the structural reason is that no public corpus of any kind was ever recorded on a phone: a recent survey's 36-row table of meeting corpora contains no smartphone entry, and the nearest neighbours are smart glasses, body-worn binaural rigs and purpose-built conference devices. The simulation shortcut is closed in both directions, because the field accepts simulated far-field audio for training and rejects it for evaluating a hardware claim, and the NOTSOFAR-1 simulator itself rests on 15,000 physically measured room responses, so the route is a parallel-capture protocol on AliMeeting's topology, an eight-channel array plus a headset on every participant across 118.75 hours in 13 rooms of 8 to 55 square metres at reverberation times of 0.3 to 0.6 seconds, with one identical recognizer decoding every stream because a different recognizer per device fatally confounds the hardware test, significance from the NIST matched-pairs sentence-segment test, and reference transcripts produced by people listening from scratch because the NOTSOFAR-1 creators found that annotators accept plausible but wrong machine guesses. What would make the asset ours is the axis nobody has, placement: face up on a table, face down, shirt pocket, trouser pocket and bag, for which this project's evidence supplies a mechanism for all five and a measurement for none. The second payoff is already sized, with off-the-shelf Whisper at 81.6 percent word error rate on the distant microphones of the LOTUSDIS corpus falling to 49.5 after fine-tuning on distance-diverse overlapping conversational data (arXiv 2509.18722), and synthetic mixtures alone giving 16.0 percent time-constrained speaker-attributed word error rate on AMI's single distant microphone and 20.1 on NOTSOFAR-1 single-channel before any real in-domain audio moved both (arXiv 2605.15442, measured with ground-truth speaker segmentation supplied), while the three reasons the field avoids phones, automatic gain control breaking the linear amplitude relationship between channels, clock drift with no shared word clock, and firmware tuned for a talker at roughly 30 centimetres to 1 metre, are platform choices with documented workarounds rather than laws of nature.

P1 , Pre-installed on the fleet Google does not reach

The free pre-installed competitor is not evenly distributed and the competitive map says so in its own Google profile: Pixel Recorder does the whole job offline with speaker labels since its version 4.2 and is Pixel-only, so it does not reach most Android users, and the developer-facing generative mode of Google's on-device machine learning kit is hardware-gated to Pixel 10 and Pixel 11, which leaves everybody else shipping Android with the demand and none of the feature. The sharper half is that the four manufacturers whose power managers kill a third-party recorder after screen lock, each behind a different multi-step settings path that a major update silently reverts, are the same four parties who could make that recorder un-killable by shipping it inside the system image. A capture path integrated below the application layer is exempt from its host's own process killer, needs no store review and no background audio declaration, and arrives pre-installed rather than through the ranking fight the competitive map calls undifferentiated crowding. The proof-of-concept shape is a capture-survival comparison on one manufacturer's handset, the same recorder run as an installed application and as a privileged system component over long sessions, passing on a survival gap large enough that the manufacturer is buying a defect fix rather than a feature.

V4 , The works council package

In Germany the purchase is gated before first use rather than litigated after it, since the consent-law research reports the Federal Data Protection Act section 26 and the Works Constitution Act section 87(1)(6) as requiring a co-determination agreement with the works council before an employer may deploy software of this kind, absent which the deployment is illegal and the council can compel the employer to block it. That blocker sits directly in front of the only compounding revenue route this project has, and it is a document problem rather than a technology problem. The product is the application plus a deployment package built to be signed rather than negotiated: a written retention window, per-participant refusal implemented as a behaviour of the software rather than asserted as a policy sentence, since GDPR Article 6(1)(a) consent must be obtainable from each participant independently and refusable without professional detriment, no model training as a contract term, no biometric enrolment as an architectural fact, and an Article 28 data processing agreement in the box. That the gate is real and closes hard is already recorded in Chapman University's August 2025 blanket ban on Read AI over data security, privacy and consent concerns, and the honest tension is that the individual buyer asks for processing on their own device and not for certifications, which is a different evidence base and a different sale.

G5 , The visit note that lands in the record

The outputs go into the customer relationship management (CRM) record rather than into a notes archive, and the organisation buys because notes typed in by the representative are the known data-quality hole in every CRM deployment. This is the one buyer where the in-person meeting somebody walks out of has money attached to the write-up, and where no competitor can reach: a bot cannot join a customer visit and a conferencing platform has no stream to hand over, which is why Claap, Zocks and FinMate all built this shape for calls and none of them for the room. The evidence that this is the same person the seed describes is that the buyer sample was drawn partly from the sales and consulting communities, the verbatim call-capture statement comes from a sales representative, and the buyer study describes the purchase path as an individual buying one device outside their organisation's technology function, proving it works, then acquiring units for colleagues. The gate is a security review, SOC 2 Type 2 certification plus a data processing agreement, stated data residency and corporate login, which is a different body of evidence from the one that convinces the person pressing record.

The pivot tier

A pivot here is an idea that contradicts one of this project's own scope rules: the product ships as software on a device the customer already owns, the phone's own microphones are the capture path, the interaction is two taps, the deliverable is a transcript, a summary and action items, the meeting is one the user physically walks out of, the outputs are ready within the few minutes after it ends, and the target is the person who buys or would buy a dedicated recorder. These ideas are kept and argued rather than smoothed away, each one named with the rule it breaks, because an idea that contradicts the thesis is evidence about the thesis.

ID Novelty Defensibility Time-to-market Strategic fit Total
M4 3.9 2.9 2.3 2.0 11.1
T6 1.3 2.7 2.1 2.0 8.1
V5 1.1 1.0 5.0 1.0 8.1
P2 2.3 3.4 1.0 1.0 7.7
S6 3.1 2.3 1.0 1.0 7.4
G6 (folds F6) 2.1 1.3 1.7 1.0 6.1

M4 , The meeting is the array

The platform will not give one phone four channels, but it will give four people one channel each, and the claim has to be precise about what that buys: independent phones share no hardware clock and drift measurably over a session, which destroys the microsecond phase alignment that beamforming needs, so guided source separation and any classical beamformer are out and must not be promised. What alignment after the fact does support is everything that needs tens of milliseconds rather than microseconds, since cross-correlation alignment is the accepted method for independent devices in the CHiME-6 challenge recipe, and once two streams are aligned at that resolution the per-turn energy ratio between phones is a direct speaker-attribution cue, while channel selection by envelope variance, a CHiME-7 baseline technique, picks whichever phone was nearest the person talking. Every microphone is still a phone's own, so the no-new-hardware and phone-is-the-microphone rules both hold and only the interaction rule breaks. One side effect is worth as much as the acoustics: each participant starting their own capture is consent expressed as an action rather than as a disclosure screen, which speaks to the twelve-jurisdiction all-party set and to the covert-capture theory in Chamberlain v. Granola, a case brought against a product that processed on the device, over how the capture was presented rather than where the computation ran.

T6 , Free at the person, paid at the company

The funding map's first emerging pattern is that every revenue route that compounds in this field ends at an organisation and not at a person, and the evidence that opens the compounding route is a security review, SOC 2 Type 2, ISO 27001, a data processing agreement and stated data residency, while the sources say the individual and small-team buyer asks for something else entirely, processing on their own device, which is two different products' worth of evidence. Set that beside the crowding pattern, nine or more vendors selling the same three outputs from the same rented models at 9 to 19 dollars, three of them giving the core interaction away and both operating systems shipping it pre-installed and free, and a consumer subscription cannot pay an acquisition cost out of what one person spends in a category whose own price floor is zero. The pivot reorders the business: the individual product becomes the acquisition channel, given away and built so that giving it away costs nothing, since on-device transcription sits near 0.05 US dollars a recorded hour and a fully local pipeline at zero, and the revenue is the seat contract sold into the companies where those free individuals turn out to cluster, which is the motion this field's winners run, with Otter growing on Business and Enterprise, Read AI valued on Fortune 500 penetration rather than on the 100,000 consumer accounts it adds weekly, and Granola's Series C raised on an enterprise-context thesis. Two honest costs: the seed's customer stops being the payer and becomes the distribution, and the certification toll is paid in full before the first seat contract is signed, in a currency the free user does not value and will not fund.

V5 , The record you speak yourself

Do not record the room at all: on the walk back the user taps start, says what happened, taps stop, and gets back a transcript of their own words, a summary of it and the action items in it. Every legal exposure in this project's evidence is defined by capturing the voice of somebody who never agreed to anything, with four pending United States actions on record (In re Otter.AI Privacy Litigation arising from Brewer v. Otter.ai Inc., Cruz v. Fireflies.AI Corp., Basich v. Microsoft Corp. and the Chamberlain action against Granola), the German Criminal Code section 201 making it a criminal offence for a participant to record the non-publicly spoken word at up to three years, French Penal Code Article 226-1 requiring all-party consent with a proportionality test from the French data protection regulator on top, and every one of the twelve in-person all-party jurisdictions regulating the capture of another person. It is also the acoustically trivial case, one close-talking speaker rather than the far-field multi-talker condition the whole far-field effort exists to survive, which is why it runs on the handset at zero marginal cost and finishes inside the walk back. Argued honestly it gives up a great deal: it serves none of the 64 stated hardware purchase reasons, it cannot touch the 28.1 percent cellular-call segment any more than the original can, and it loses the verbatim record that is the thing a meeting recorder is bought as insurance for.

P2 , Call recap licensed into the carrier's dialer

The funding map closes third-party call capture as structurally dead on both platforms, iOS exposing no public interface to the audio of an active call and Google Play banning the accessibility-service route since May 2022, and it names the one entity for whom it is not closed, whoever owns the dialer. The competitive map's SK Telecom profile is the working instance, an assistant built into the native dialer that transcribes, translates and summarises for over ten million subscribers with no download and no app store, described there as the shape of the threat no application can answer, because it does not compete for the download. Against that sits the buyer evidence: 18 of 64 explicit purchase-rationale statements, 28.1 percent and the largest single category, are about capturing a cellular telephone call, quoted verbatim as "Otter doesn't record phone calls (thanks, iOS)", and those buyers did not walk past an equivalent application, because there is none and can be none. The proof-of-concept shape is a single carrier's answer to whether it can originate call recording for its own subscribers in its own market, where pass is a yes with a named legal basis and kill is anything softer.

S6 , Phone-condition recognition sold by the hour

This rail competes on price per minute against companies whose prices are already published, Deepgram Nova-3 at about 0.258 US dollars per audio hour, AssemblyAI at a 0.15 base and recorded at 0.37 with speaker separation and entity detection switched on, Groq quoted near 0.04, so a general recognizer sold this way is a commodity and we would lose on price. The only defensible wedge is the condition rather than the model: none of those vendors has ever been measured on a phone in a room, their published accuracy figures name neither a corpus nor a microphone condition, and a supplier that can state a time-constrained speaker-attributed word error rate on phone-captured far-field meeting audio at a named placement against a named reference sells the one thing nobody else in the supply chain can state. The buyers who would care are the ones who cannot reach the condition themselves, the wearable and card-recorder vendors whose devices sit in the same acoustic position and the application developers whose products already record through the phone. Selling inference is the version that survives, because licensing the recorded audio itself is closed by the voiceprint statutes and by the category's own contractual promises never to train on customer data.

G6 (folds F6) , The accessory we do not build

The buyer evidence is the whole argument: all 64 of 64 stated purchase reasons sit in categories the buyer study classifies as either structurally impossible for a third-party application or inherent to owning a separate object, call capture at 28.1 percent, battery preservation at 14.0, physical start-stop tactility at 7.8, and not one buyer mentions audio quality. The field has already run this experiment and it ran the other way round from the hardware incumbents: ByteDance kept Lark as the product and let Anker build the recording bean as an accessory that syncs into the software, hardware treated as an accessory to software rather than as the product, so the software keeps its economics and its rails, the object is somebody else's inventory, warranty and returns, and the customer gets the separate battery and the physical button that 21.8 percent of stated reasons are about. The counter-evidence is stated plainly rather than argued around: this is the second-object pattern that killed Humane and left Limitless owners holding a brick when Meta switched the service off, the competitive map rates that pattern as the one thing this project is currently immune to, and the pivot contradicts the seed's entire ten-times-better claim. Folded in from the founder seat is the passive variant of the object, a start and stop button and nothing else, no microphone, no data connection and no audio path, which turns entirely on whether any documented interface lets a paired accessory start third-party microphone capture on a locked phone.

What was dropped, and why

Forty-one ideas were proposed and twenty-seven entries were scored. Nothing was erased; every idea that is not a scored entry above was either folded into one that is, or dropped for a stated reason, and both are listed here.

What this slate does not answer

A reader of the pitch arrives holding one line, that we give the person who was about to buy a recording device the recording device they already own, and leaves this slate without a single idea that tests whether the phone can do the job. The open question of what word error rate a real phone achieves in a real meeting at five named placements, face up on a table, face down, shirt pocket, trouser pocket and bag, is measured by nothing here; the two entries that propose building the instrument to measure it are the two the audit tagged as having no payer. Nor does the slate engage this project's own first kill criterion, which the evidence says fires on its literal wording: the product already exists free, Voicenotes does the exact flow at zero, Fathom and Granola and Hedy give the interaction away, and both operating systems ship it pre-installed. Almost every idea here is a reason to prefer our recorder over another application. Two are reasons to prefer an object, and they answer the acquisition question by putting the 159 dollars back. None is a reason for the person holding the device to put it down, which is the sentence the pitch is made of.

The cheapest observation that would most change this ranking is the interruption protocol from one of the folded capture-survival entries: take two shipping recorder applications, one iPhone and one Samsung handset, force four conditions on each, an incoming call mid-session, a screen locked for an hour, the manufacturer's own battery daemon, and a second application requesting the microphone, and count how many sessions come back with complete audio. It needs no corpus, no annotation and no counsel, and it is the only test on the slate whose outcome moves six ideas at once in either direction. If the incumbents survive it cleanly, the highest-rated cluster in the round has nothing to sell against and the slate collapses to the announcement group and the employer-seat group. If they fail it, that cluster becomes the only mechanism in this project that is simultaneously buildable, unclaimed and evidenced, and every other idea should be re-ranked beneath it. Two other observations are comparably cheap but move less: reading Apple's and Google's published policy for their zero-install surfaces settles the record-before-install entry in one paragraph and touches one idea, and opening one named flagship handset to see whether its built-in recorder already ships all three outputs settles the pre-installed entry and the unverified claim about Samsung's own recorder, and touches two.

View as PDF Download PDF