Preprint / Version 1

How Does the Optimal Number of Fourier Series Terms Required to Reconstruct an Alexa Voice Command Vary Between Different Speakers?

##article.authors##

  • Aarush Raheja Symbiosis Internationa School Pune

DOI:

https://doi.org/10.58445/rars.4111

Keywords:

fourier series, speech reconstruction, voice signals, signal processing, root mean square error, spectral energy, alexa

Abstract

I investigated how the number of Fourier series terms needed to reconstruct the same Alexa voice command changes between different speakers. I recorded the command “Alexa, what’s the weather like?” using a softer voice and a louder, deeper voice under the same conditions. Each recording was converted into numerical data and reconstructed using finite Fourier series with different numbers of terms. I measured the accuracy of each reconstruction using root mean square error and the proportion of Fourier energy captured. I defined the optimal number of terms as the smallest value that captured at least 95% of the reference energy while producing less than a 1% improvement in RMSE when another term was added. The softer voice required 193 Fourier terms, while the louder voice required 194 terms. This shows that the number of terms needed can vary between voice samples, although the difference found in this experiment was small. The results also show how Fourier series can be used to represent differences in the frequency content of human speech.

References

Innovation World. (2025). Parseval’s theorem. https://innovation.world/invention/parseval-theorem/

Weisstein, E. W. (2008). Fourier series. MathWorld, Wolfram Research. https://mathworld.wolfram.com/FourierSeries.html

Downloads

Posted

2026-08-30