Coherence Research

Working paperNot peer-reviewed

ContentIQ Working PaperCIQ-WP-2026-16Version 1

Talking to camera sells more per view than trending sounds on TikTok Shop

13,304 videos tracked, 3,342 analysed in depth, 29 categories, 2 Sept–1 Oct 2026

Coherence Research

Scope
Topic
Data
2 Sept – 1 Oct 2026
Sample
13,304 videos · 3,342 watched · 29 categories

Key findings

  1. 1.14×sales per view for speaking on camera (90% range 1.08–1.19), the highest of the main audio types
  2. 0.76×sales per view for trending sounds (0.70–0.81), despite 13.5% of watched-video sales
  3. 1.08×sales per view for enthusiastic tone (1.04–1.12), which carries 66.5% of sales
  4. 0.50×sales per view for emotional tone (0.43–0.58), the lowest of the tones
  5. 1.67×sales per view for urgent tone (1.28–2.17), but from only 32 videos
  6. 0.64AUC of the top-seller model (0.57–0.70): useful, but far from decisive

Abstract

We examine what top-selling TikTok Shop videos sound like and which audio and tone choices are associated with more sales per view. Speaking on camera (1.14× the market, 90% range 1.08–1.19) and voice-over (1.05×, 1.01–1.10) sell above average per view, while trending sounds (0.76×, 0.70–0.81) and music-only videos (0.85×, 0.77–0.93) sell below it. Among tones, enthusiastic delivery carries two thirds of sales at 1.08×, while emotional delivery sells at 0.50× the market. A model of top-selling videos has modest predictive power (AUC 0.64, 0.57–0.70), so audio and tone are one input among many, and all findings are correlations from a one-month window.

1. Introduction

Brands and creators on TikTok Shop make a series of choices about sound before they think about anything else: whether to speak to camera, record a voice-over, use a trending sound or let music carry the video, and what tone to take. These choices are cheap to change, so it matters whether they are associated with sales. This paper asks which audio and tone choices appear in top-selling videos and which sell more per view.

Prior work on short-video popularity shows that attention is long-tailed and that early popularity predicts later popularity [1]. State-of-the-art popularity prediction finds that creator features matter most [2], and multimodal benchmarks combine audio, visual and text signals [3]. Those studies predict views or popularity. We look at sales per view in a shopping setting, where the audio is part of a sales pitch.

We use correlation language throughout. Returns to marketing are hard to measure from observational data [4], and nothing here shows that changing a video's sound would change its sales.

2. Data

Our analysis of TikTok Shop tracked 13,304 videos between 2 September and 1 October 2026, across 29 categories. Our AI analyst watched and analysed 3,342 of these in depth. This is a count of videos, not of views.

For each analysed video we recorded the audio type (speaking on camera, voice-over, music only, trending sound or product sounds) and the tone of delivery (for example enthusiastic, calm, funny, emotional, urgent). Shares of sales and sales per view rest on these analysed videos.

The window is a single month, so seasonal effects, including retail moments, are not separated out. Some groups are small and are flagged where they appear.

Daily rankings of top-selling TikTok Shop videos in the US were collected from several independent sources and merged into one record per video. Where sources overlap, their sales estimates are cross-checked against each other, and videos whose estimates disagree by more than half are flagged. An AI analyst watched each video and recorded its opening, format, angle, production style, use of AI, hook source, length, pacing and what appears in the first second, without seeing how the video sold. Sales per view is a video group’s revenue per view divided by the market’s, shrunk toward the average for small groups, with 90% ranges. This paper covers 2026-09-02 to 2026-10-01: 13,304 videos tracked, 3,342 of them watched and analysed. Predictions come from a logistic model retrained daily and tested on recent weeks it had not seen.

3. Methods

Sales per view compares a group of videos' revenue per view with the whole market's, so 1 is the market average. Estimates for small groups are shrunk toward the average by empirical Bayes [5]. We report 90% intervals, computed on the log scale.

Share of sales is the group's portion of the revenue of the watched videos. A group can have a large share of sales simply because it has many videos, so share and sales per view answer different questions.

The predictive model is a logistic model of top-selling videos, retrained daily and tested on recent weeks it had not seen (a time split). We report AUC as a ranking measure [6], with a bootstrap range [7].

4. Results: audio type

Speaking on camera sells at 1.14× the market per view (90% range 1.08–1.19) across 1,067 videos and holds 34.3% of sales. Voice-over is the most common choice, with 1,418 videos and 45.1% of sales, and sells at 1.05× (1.01–1.10). Both ranges sit above 1, though voice-over only just.

Trending sounds sell at 0.76× (0.70–0.81) with 450 videos and 13.5% of sales. Music-only videos sell at 0.85× (0.77–0.93) with 247 videos and 6.4% of sales. Both ranges sit clearly below 1. Product sounds sell at 0.53× (0.40–0.70), but only 29 videos support this and confidence is medium, so we treat it as indicative.

Videos where a person talks, to camera or in voice-over, sell more per view than videos led by a sound or music. Trending sounds still account for a sizeable share of sales because they are used often, not because they sell more per view.

Table 1. Sales per view and share of sales by audio.
ValueVideos (n)Share of salesSales per view90% range
Speaking on camera1,06734.3%1.14×1.08–1.19
Voice-over1,41845.1%1.05×1.01–1.10
Music only2476.4%0.85×0.77–0.93
Trending sound45013.5%0.76×0.70–0.81
Product sounds290.7%0.53×0.40–0.70

Note. Sales per view is relative to the market (1.00 = average), shrunk toward the average for small groups (empirical Bayes); ranges are 90% intervals on the log scale. Source: ContentIQ analysis of 13,304 TikTok Shop videos, 2 Sept – 1 Oct 2026.

5. Results: tone of delivery

Enthusiastic delivery dominates, with 2,021 videos and 66.5% of sales, and sells at 1.08× (1.04–1.12). Informative delivery holds 14.3% of sales and sells at 1.03× (0.96–1.11), indistinguishable from the market average.

Calm (0.83×, 0.77–0.89), funny (0.75×, 0.66–0.84) and emotional (0.50×, 0.43–0.58) tones sell below the market per view. Emotional is the weakest tone, with 99 videos and 2.5% of sales.

Urgent (1.67×, 1.28–2.17) and sceptic-convinced (1.32×, 1.07–1.62) tones sell highest per view, but they rest on only 32 and 47 videos and together hold 2.8% of sales. Their ranges are wide, so these are leads to test rather than established effects.

Table 2. Sales per view and share of sales by tone.
ValueVideos (n)Share of salesSales per view90% range
Urgent320.8%1.67×1.28–2.17
Sceptic convinced472%1.32×1.07–1.62
Enthusiastic2,02166.5%1.08×1.04–1.12
Informative49314.3%1.03×0.96–1.11
Calm40410.7%0.83×0.77–0.89
Funny1163.1%0.75×0.66–0.84
Emotional992.5%0.50×0.43–0.58

Note. Sales per view is relative to the market (1.00 = average), shrunk toward the average for small groups (empirical Bayes); ranges are 90% intervals on the log scale. Source: ContentIQ analysis of 13,304 TikTok Shop videos, 2 Sept – 1 Oct 2026.

6. Results: how much the model can say

The top-seller model reaches an AUC of 0.64 (90% range 0.57–0.70), trained on 2,667 rows. It was not trained only on analyses that never saw sales (cleanOnly is false), so its figure may be somewhat flattering and should not be read as a clean forecast.

The largest drivers were not audio or tone. Filmed live action had the highest odds ratio (2.54), followed by a visual pattern-break combined with a review or testimonial (1.92) and no visible AI (1.71). Before-and-after formats (1.66) and boosting with ads (1.59) were also associated with top-selling status, while showing the product in the first second (0.62) and education content (0.59) were associated with lower odds.

An AUC in this range means the model ranks top sellers better than chance but misses many. Audio and tone should be read as one part of a larger picture.

7. Discussion

For brands and creators, the evidence favours a human voice. Speaking on camera and voice-over both sell at or above the market per view, while trending sounds and music-only videos sell below it. Enthusiastic delivery is the safe default, since it combines high volume with above-average sales per view. Calm, funny and emotional tones sell below average per view and may need a strong reason to be chosen. Urgent and sceptic-convinced tones merit controlled tests.

This fits prior work that finds creator and content features dominate popularity [2], though those studies measure popularity, not sales. A plausible alternative explanation is that categories, creators and products differ across audio and tone groups. For example, a product that needs explaining may favour speech, and trending sounds may be paired with products that are less suited to direct selling. Ad boosting, which the model associates with top-selling status, may also differ between groups.

The trending-sound result may reflect sales per view rather than reach: such videos may gather many views from viewers who are not in a buying mood. We cannot separate these explanations, and advertising effects are hard to identify from observational data [4].

8. Limitations

Most videos in the sample are top-ranked, so findings mostly separate strong sellers from good ones rather than from all videos; typical and weak videos are being added. Results are associations, not causes. Paid promotion, creator audience size and product price are not fully controlled for. Revenue, views and sales are estimates, not figures reported by TikTok. Urgent (32 videos), sceptic-convinced (47) and product sounds (29) rest on small samples, and product sounds has medium confidence. Results come from a single 30-day window and pool all 29 categories, so category mix may drive some differences. The model was not trained only on analyses that never saw sales, and its odds ratios are associations, not effects.

9. Conclusion

In this month of TikTok Shop data, videos in which someone speaks sell more per view than videos led by trending sounds or music, and enthusiastic delivery is the most common and an above-average tone. Emotional and funny tones are associated with lower sales per view.

These are associations from one 30-day window, several groups are small, and the predictive model is modest (AUC 0.64, 0.57–0.70). We recommend treating the findings as hypotheses for brands to test within their own category and format, and will update them as the window extends.

References

  1. [1]Szabo, G., & Huberman, B. A. (2010). Predicting the popularity of online content. Communications of the ACM, 53(8), 80–88. cacm.acm.org/research/predicting-the-popularity-of-online-content
  2. [2]Ye, L., Zhang, Y., Wu, Y., et al. (2025). MVP: Winning solution to SMP Challenge 2025 video track. arXiv:2507.00950. arxiv.org/abs/2507.00950
  3. [3]Lu, J., Wang, W., Xiao, M., et al. (2024). M3TR: Temporal retrieval enhanced multi-modal micro-video popularity prediction. arXiv:2411.15455. arxiv.org/abs/2411.15455
  4. [4]Lewis, R. A., & Rao, J. M. (2015). The unfavorable economics of measuring the returns to advertising. The Quarterly Journal of Economics, 130(4), 1941–1973.
  5. [5]Efron, B., & Morris, C. (1975). Data analysis using Stein’s estimator and its generalizations. Journal of the American Statistical Association, 70(350), 311–319.
  6. [6]Fawcett, T. (2006). An introduction to ROC analysis. Pattern Recognition Letters, 27(8), 861–874.
  7. [7]Efron, B., & Tibshirani, R. J. (1993). An Introduction to the Bootstrap. Chapman & Hall.

Cite as

Coherence Research (2026). Talking to camera sells more per view than trending sounds on TikTok Shop. ContentIQ Working Paper CIQ-WP-2026-16, version 1. https://www.coherenceltd.com/research/audio-and-tone

Version history

  1. Version 1This version