The Commonplace
Home Three-study pilot Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Automated voices already power a sizeable slice of unwanted calls: a U.S. honeypot finds roughly 27% open with machine audio—split between replays and detected synthetic speech—yet callers almost never admit they're automated (0.44%), and detector labels require listener validation.

The Machines Are Calling: Measuring Automated and Synthetic Voices in Unwanted Inbound Calls
Xingyu Shen, Tommy Duong, Muduo Xu, Xiaodong An, Jiaqi Gan, Haoyuan Tang, Jamey Z. Liang, Siyu Zhang, Yan Zhang, Simiao Ren · September 10, 2026
arxiv descriptive medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Xingyu Shen unresolved corpus identity
  2. Tommy Duong unresolved corpus identity
  3. Muduo Xu unresolved corpus identity
  4. Xiaodong An unresolved corpus identity
  5. Jiaqi Gan unresolved corpus identity
  6. Haoyuan Tang unresolved corpus identity
  7. Jamey Z. Liang unresolved corpus identity
  8. Siyu Zhang unresolved corpus identity
  9. Yan Zhang unresolved corpus identity
  10. Simiao Ren unresolved corpus identity
A U.S. lead-generation honeypot shows at least 26.9% of unwanted inbound calls open with machine-produced audio (about 13.8% replayed recordings and 13.1% detector-labeled fresh synthesis), callers almost never disclose automation (0.44%), and blinded listeners confirm roughly 54% of detector-flagged clips.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

In February 2024 the U.S. Federal Communications Commission (FCC) placed AI-generated voices under the Telephone Consumer Protection Act (TCPA). Yet no peer-reviewed measurement says how much unwanted call traffic is placed by a machine, or how much of that machine speech is synthesized rather than played from a recording. We report both with a disclosed pipeline. An interactive voice honeypot (language-model personas on real U.S. numbers, the caller recorded on its own track) recorded 10,987 calls over 66 days; 11 days on which our stack answered silently are set aside. Three instruments read each opening: an audio fingerprint that finds the same recording played on other calls, a commercial synthetic-speech detector on the caller's first ten seconds, and blinded listeners who check what it flags. Of the 7,233 calls our persona greeted on normal days, 13.8% open with a recording we also heard on another call, and 13.1% with fresh audio the detector labels synthetic. A further 9.9% open with a caller who never spoke after our greeting, 54.2% with fresh audio the detector labels human, and 9.0% could not be scored. Machine-voiced openings are therefore at least 26.9%, a further tenth of calls are silent connections we read as machine-placed, and replays of a recording make up 45% of the detector's own rate (29.3% of 6,192 scored openings). The same waveform played on two calls lands on opposite sides of the detector's threshold 13.6% of the time, and eleven listeners confirm 54.4% of what it flags. Synthetic openings concentrate in lead-generation spam (33.8%), not fraud (21.1%); 0.44% disclose automation. Prevalence tracks how long a bait number has circulated (59% against 19% in the same weeks): seeding history, not calendar time, explains the trend. Campaigns outlast their numbers: one recorded compliance notice opens calls in six campaigns, and one synthetic voice serves nine.

Summary

Main Finding

An interactive voice honeypot collected 10,987 inbound unwanted-call recordings (66 days). Of the 7,233 calls the honeypot greeted on normal (non-outage) days, at least 26.9% opened with machine evidence on the audio alone: 13.8% opened with a recording that reappeared on another call and 13.1% with fresh audio that a commercial synthetic-speech detector labeled synthetic. A further 9.9% were silent connections (no reply after the persona’s greeting) that the authors read as machine-placed; 9.0% of openings could not be scored. The commercial detector by itself labeled 29.3% of 6,192 scored openings synthetic; blinded listeners confirmed 54.4% of the detector-flagged clips, giving a listener-confirmed synthetic-opening estimate of roughly 15.6–15.9% of scored calls under plausible reweighting assumptions. Synthetic openings concentrate in lead-generation spam (33.8%) more than in fraud (21.1%), and only 0.44% of synthetic-labeled calls verbally disclosed that they were automated.

Key Points

  • Corpus and scale: 10,987 calls recorded over 66 days; main analytic stratum is 6,619 two-turn calls, of which 6,192 carried scorable caller speech.
  • Measurement instruments:
    • Audio fingerprint (detector-free) to identify identical replayed recordings across calls.
    • Commercial synthetic-speech detector applied to caller-only 10-second onset-aligned clips (per-utterance scores aggregated by max; threshold dfmax ≥ 0.85 used as primary operating point).
    • Blinded human listeners (11 internal listeners) for validation.
  • Main numeric results (normal-day greeted calls, unless noted):
    • 13.8% replayed recordings (fingerprint).
    • 13.1% fresh audio labeled synthetic by detector (29.3% of 6,192 scored openings when using detector alone).
    • Combined machine-voiced (recording + detector-labeled synthetic) = 26.9% (audio-evidence only).
    • 9.9% silent connections after greeting (read as machine-placed).
    • 9.0% unscorable openings.
    • Blinded validation: of 1,816 detector-flagged calls, 878 were judged by listeners and 54.4% of those confirmed synthetic; only 23 judgments exist for detector-nonflagged clips in this report (so authors present precision, not a fully validated prevalence).
    • Same-waveform instability: a waveform played on two calls crosses the detector threshold oppositely ~13.6% of the time.
  • Distributional facts:
    • Synthetic-labeled calls more common in lead-generation sales (33.8%) than fraud (21.1%).
    • Synthetic-labeled calls often use toll-free numbers and tend to be shorter; differences in conversion or credential requests largely vanish after adjusting for call length and number type.
    • Disclosure is rare: 0.44% admit present automation; 17% announce that the call is recorded.
  • Campaign structure and asset reuse:
    • Campaigns often reuse the same script/recording or the same synthetic voice across many disposable originating numbers. Example: one recorded compliance notice opened calls in six campaigns; one synthetic voice served nine campaigns.
    • Seeding history (how long a bait number has been exposed in lead forms) predicts synthetic prevalence better than calendar time; exposure-age confounds naive trend analyses.
  • Sensitivities and thresholds:
    • Fingerprint threshold choice affects split between “replay” and “fresh synthetic” (e.g., stricter fingerprint → replay 9.2%, synthetic 16.7%; total machine-voiced changes ≈1 percentage point).
    • Detector threshold choice would materially change point estimates (authors note the score distribution trough near 0.45–0.50; placing a threshold there would yield a higher synthetic rate).

Data & Methods

  • Honeypot design:
    • Interactive voice honeypot: conversational LLM personas, each on a dedicated real U.S. telephone number, answer inbound calls and converse; persona audio and caller audio are recorded on separate channels so the caller leg is isolated for scoring.
    • Numbers were seeded into U.S. lead-generation web forms using fabricated consumer records; in forms with a consent checkbox the seeding agent sometimes checked it (seeding history is recorded and used in analyses).
  • Corpus and filtering:
    • Collection span: 28 May–21 July 2026 (66 days). Total recorded calls with caller track: 10,987.
    • Two-turn analytic stratum: 6,619 calls (both honeypot greeting + at least one transcribed item); 6,192 carried scorable caller speech (rest unscorable for technical reasons).
    • Eleven outage days (persona answered silently) are set aside per a stated rule; calls during outages were handled separately.
  • Audio preprocessing:
    • Caller-only audio onset detection: energy-based detector (30 ms frames, 10 ms hop), onset declared when energy > noise floor +10 dB for 250 ms; 200 ms pre-roll retained.
    • Fixed 10-second window from detected onset used for detection/validation (chosen for antispoofing literature and detector performance).
  • Replay detection (audio fingerprint):
    • Vendor-free audio fingerprinting finds identical waveforms played across calls; used to label “replay” recordings versus fresh renderings from a synthesizer.
    • Fingerprint has its own threshold; authors report results under two settings to show sensitivity.
  • Synthetic-speech detection:
    • Commercial closed detector (vendor withheld) with per-utterance deepfake_score ∈ [0,1]; authors aggregate to dfmax = max per-utterance score per call and use dfmax ≥ 0.85 as primary label for “synthetic.”
    • Authors also report mean-aggregation for sensitivity and sweep thresholds.
  • Human validation:
    • 11 internal listeners, blinded interface, labeled ten-second clips; no external screening or qualification reported.
    • Validation sample: 1,816 detector-flagged calls, 878 judged by listeners (54.4% confirmation); only 23 judgments for non-flagged snippets in this version, so validation supports detector precision estimate rather than full calibrated prevalence.
  • Limitations discussed by authors:
    • Closed commercial detector (no training-data transparency) applied to narrowband 8 kHz telephony audio—benchmarks can degrade out-of-distribution.
    • Listener panel internal, unscreened—may bias validation.
    • Threshold choices (detector and fingerprint) materially affect point estimates; authors provide sensitivity sweeps.
    • Honeypot seeding schedule and fabricated identities mean prevalence estimates are conditional on the collection design (exposure-age effects).
    • Some calls unscorable or silent; outages handled per rule.

Implications for AI Economics

  • Adoption and substitution effects
    • Substantial prevalence (≈15–27% depending on instrument and validation assumptions) implies voice synthesis is a material delivery mechanism in unwanted inbound calls already by mid‑2026. This supports a view that AI voice is not merely marginal novelty but an operational input at scale in high-volume lead-generation.
    • Lower per-call marginal cost of synthesized or automated voice (vs. live agents or paid voice actors) makes scaling cheap outreach feasible; firms in low-harm, high-volume segments (lead gen/sales) have clear incentive to adopt it.
  • Assetization and firm strategy
    • Campaign assets (scripts, recordings, synthetic voices) persist across disposable numbers. This means the durable economic value to an operator is in voices/scripts/models rather than numbers. Markets will price and trade these assets (voice profiles, script templates, reusable recordings), increasing secondary markets for voice assets and monetization opportunities for TTS/voice cloning providers.
    • Blocklisting/reputation systems that index by originating number face a structural limitation because numbers are cheap and disposable; economic mitigation will shift toward content-/asset-based signals (script fingerprints, voice-model identifiers) and upstream carrier/platform interventions.
  • Regulation, compliance costs, and market structure
    • Low voluntary disclosure (0.44%) implies a large gap between existing practice and any disclosure mandate (e.g., proposed rule requiring callers to state that the voice is AI-generated). If regulators impose disclosure/identification/opt-out obligations (as the FCC’s TCPA regime contemplates), firms using synthetic voice will face compliance costs (disclosure logic in calling stacks, provenance attestation, audit trails), and some operators may exit or move to more adversarial tactics to avoid compliance.
    • Enforcement and compliance monitoring will be a service market—opportunities for vendors offering verified provenance, auditable attestation, or network-level detection/provenance stamping.
  • Labor market and task allocation
    • The concentration of synthetic voice in lead-gen (rather than high-harm fraud) suggests initial displacement pressure on lower-skilled outbound sales agents and call-center roles. Firms may reallocate human agents toward higher-complexity or closer stages (sales closers), while automated front ends handle scale outreach.
    • Authors found AI-labeled calls often hand off to live closers; however, they did not find strong evidence that synthetic openings convert less once length is controlled. That affects firms’ ROI calculations for automation vs. human staffing.
  • Detection and mitigation economics
    • Detection noise (same waveform can cross detector threshold inconsistently ~13.6% of the time) and limited validation precision (~54% confirmed) indicate nontrivial enforcement and false-positive costs. Deployers of detection (carriers, handset vendors) must weigh the economic costs of false positives (blocking legitimate calls) versus benefits of blocking synthetic nuisance traffic.
    • The need for improved, robust telephony-specific detectors creates commercial opportunity and justifies investment in research/benchmarking for narrowband telephone conditions.
  • Market externalities and social cost
    • Widespread use of synthetic voices in nuisance calls can reduce consumer trust in voice contact channels, lowering the value of phone-based outreach for legitimate businesses (a negative externality). This raises potential justification for regulation or platform-level mitigation that internalizes externalities (e.g., stricter provenance requirements, liability rules).
  • Policy design insights
    • Because persistent campaign assets (scripts/voices) matter more than originating numbers, policy and enforcement that target only numbers will lag operationally and economically; cost-effective regulation should consider content provenance, tamper-evident attestation of automation, and obligations on upstream intermediaries who provision voice-generation services.
    • Baseline prevalence reported here informs cost–benefit calculations for proposed disclosure rules and the expected scale of compliance enforcement.
  • Takeaway for economists and policymakers
    • The paper provides a replicable, instrumented baseline (with disclosed pipeline and sensitivity analyses) showing that synthetic voice is a significant operational input in unwanted calling. Economic models of scam/telemarketing markets should incorporate low marginal cost of synthetic voice, asset reuse across campaigns, disposable-number dynamics, and the nontrivial measurement uncertainty in detection when forecasting impacts of regulation or designing incentives for provenance solutions.

Limitations worth bearing in mind when using these results in economic models: the estimates are conditional on the honeypot’s seeding design and collector schedule; the detector used is closed-source and applied to narrowband telephony; human validation is partial and internal; threshold choices and fingerprint settings materially affect point estimates.

Assessment

Paper Typedescriptive Evidence Strengthmedium — Large, novel field corpus with careful instrumentation (separate caller channel, audio fingerprinting, detector scores and blinded listener checks) provides direct measurement of voice provenance; however results are constrained by sample selection (lead-generation honeypot), reliance on a closed-source commercial detector, limited and internal listener validation (small negative-sample labeling), and a short, staggered seeding window, all of which limit how conclusively the numbers generalize. Methods Rigormedium — Methodological strengths include separate-channel recording, onset-aligned 10s windows, a detector-free audio-fingerprint to identify replays, and disclosure of thresholds and aggregation rules; weaknesses include use of an opaque commercial detector without full calibration to narrowband telephone audio, internal listeners with no reported QC or interrater statistics, limited labeled negatives for prevalence calibration, and potential biases from the honeypot seeding strategy. SampleA conversational voice honeypot recorded 10,987 inbound calls (over 66 days) to U.S. telephone numbers seeded into lead-generation forms; after filtering, a two-turn stratum contains 6,619 calls, of which 6,192 carried scorable caller openings (caller and persona recorded on separate channels). Validation used a commercial synthetic-speech detector applied to 10-second onset-aligned clips (max-utterance aggregation, threshold 0.85) and blinded judgments from 11 internal listeners on subsets (878 confirmed judgments reported for detector-flagged clips; only 23 judgments for non-flagged clips in this version). Seeding was staggered (one long-lived number, then ten numbers added on 1 July), and collection ran from late May to July 2026. Themesadoption governance GeneralizabilitySample is a honeypot seeded into U.S. lead-generation funnels and may not represent all unwanted call traffic (e.g., government/telecom robocalls, international campaigns, or contact-center use cases)., Short collection window (May–July 2026) and staggered seeding schedule mean prevalence estimates depend on exposure-age dynamics of the seeded numbers., Results limited to narrowband (8 kHz) telephony audio; detector performance and fingerprinting may differ on higher-quality channels., Use of a closed-source commercial detector prevents independent verification of classification errors and training-domain mismatch., Blinded validation relied on internal listeners with no reported screening, inter-rater reliability, or attention checks, and the non-flagged sample is too small to robustly estimate false negatives.

Claims (11)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Among 7,233 calls greeted by the honeypot on normal days, 26.9% opened with machine-voiced audio: 13.8% used a recording heard on another call and 13.1% used fresh audio labeled synthetic by a commercial detector. Automation Exposure positive Share of unwanted inbound calls with machine-voiced openings
Reading fidelity high
Study strength medium
n=7233
26.9% machine-voiced; 13.8% replay; 13.1% detector-labeled synthetic
0.18
A further 9.9% of greeted calls were silent connections in which the caller never spoke after the honeypot's greeting, and the authors interpret these as machine-placed connections. Automation Exposure positive Share of inbound calls with no caller speech after the greeting
Reading fidelity high
Study strength low
n=7233
9.9%
0.09
The commercial detector labeled 29.3% of 6,192 scorable call openings as synthetic, and 45% of those detector-labeled openings were replays of recordings. Automation Exposure positive Detector-labeled synthetic-speech rate and replay share
Reading fidelity high
Study strength medium
n=6192
29.3% labeled synthetic; 45% of detector-labeled openings were replays
0.18
Blinded listeners confirmed 54.4% of the synthetic-speech flags they evaluated. Ai Safety And Ethics positive Listener-confirmed precision of synthetic-speech detector flags
Reading fidelity high
Study strength low
n=878
54.4% confirmed
0.09
The same waveform played on two calls crossed the detector's synthetic-speech threshold on opposite sides in 13.6% of paired cases. Ai Safety And Ethics mixed Synthetic-speech detector consistency across repeated playback
Reading fidelity high
Study strength medium
13.6% of pairs crossed the threshold inconsistently
0.18
Detector-labeled synthetic openings were more common in lead-generation sales calls than in calls classified as fraud: 33.8% versus 21.1%. Automation Exposure positive Synthetic-speech prevalence by unwanted-call type
Reading fidelity high
Study strength medium
33.8% in lead-generation sales calls versus 21.1% in fraud calls
0.18
Only 0.44% of calls labeled synthetic included a disclosure that the call was automated; 99.6% did not disclose automation. Regulatory Compliance negative Rate of automation disclosure in synthetic-voice calls
Reading fidelity high
Study strength medium
0.44% disclosed automation; 99.6% did not
0.18
Synthetic-voice prevalence tracked the length of time a bait number had circulated: the paper reports 59% for older exposure versus 19% for numbers seeded in the same calendar weeks. Automation Exposure positive Synthetic-voice prevalence as a function of bait-number exposure age
Reading fidelity high
Study strength medium
59% versus 19%
0.18
Calling campaigns reused assets across campaigns: one recorded compliance notice opened calls in six campaigns, and one synthetic voice served as many as nine campaigns. Market Structure positive Cross-campaign reuse of voice and audio assets
Reading fidelity high
Study strength medium
one recording in 6 campaigns; one synthetic voice in 9 campaigns
0.18
AI-labeled calls did not convert fewer targets once call length was accounted for. Consumer Welfare null_result Target conversion rate
Reading fidelity high
Study strength medium
not reported
0.18
Campaigns commonly changed originating numbers while retaining the same opening line; in the sharpest example, 68 calls came from 68 different numbers over 31 days. Market Structure positive Caller-number churn within a calling campaign
Reading fidelity high
Study strength medium
n=68
68 different numbers over 31 days
0.18

Notes