🪟 Windows TippsThe Gemini desktop app is now available for Windows(11.09.2026 um 17:06 Uhr)
🕵️ SicherheitslückenBurn Out, Or Fade Away(14.09.2026 um 14:25 Uhr)
🪟 Windows TippsProfi-Edelstahlpfanne von WMF jetzt zum halben Preis erhältlich(15.09.2026 um 08:05 Uhr)
🪟 Windows TippsThe Gemini desktop app is now available for Windows(11.09.2026 um 17:06 Uhr)
🕵️ SicherheitslückenBurn Out, Or Fade Away(14.09.2026 um 14:25 Uhr)
🪟 Windows TippsProfi-Edelstahlpfanne von WMF jetzt zum halben Preis erhältlich(15.09.2026 um 08:05 Uhr)

🔧 Programmierung 🕛 vor 1 Jahr 8 Min Lesezeit
0

SpeechTrack

↗ Quelle (dev.to)
🗣️ Stimme:
📑 Inhaltsübersicht

This is a submission for the

Ever wondered how your speech sounds to others? Clarity and speed are the heartbeat of effective communication, but striking the perfect balance can be a challenge.



Enter SpeechTrack, your real-time speech companion.

With intuitive visual indicators for Clarity and Tempo, SpeechTrack empowers you to refine your delivery through post speech Feedback.



Whether you’re preparing for a big presentation, a podcast, or an important conversation,



SpeechTrack helps you speak with confidence, precision, and impact.



Try it now and transform the way you communicate!




Technical Overview


  1. App - built using REMIX.run & React web framework tailwind.css and daisyUI for styling

  2. App modules (at a high level)


    1. Microphone - to stream Audio

    2. Websocket module for receiving messages from AssemblyAI

    3. AudioProcessor - to combine streamed Audio chunks to a .wav File object

    4. Various React components to process & assemble received messages and maintain application state

    5. Backend for front to deal with:


      • API keys and temporary tokens

      • Uploading audio,

      • Getting transcripts,

      • asking LeMUR to do evaluation of Speech



    6. Finally present the Report in MARKDOWN format

    7. Maintain a history of last 3 Speeches in localStorage on client side













Demo





Here are some use cases:







1. Learning from good speakers.



A good speaker can keep the audience spell bound. What is it that they do well that we can replicate. This inspiring speech is from the movie Coach Carter.








Here is the Speech Evaluation Report from LeMUR




Analysis & Feedback



Here's my analysis and feedback for the speech transcript in a human-readable Markdown format:





Speech Evaluation





Strengths




  • Effective Use of Pauses: With 8.75 pauses per minute (ppm), the speaker is within the ideal range of 5-10 ppm. This allows listeners to process key ideas and adds emphasis to important points.


  • Inspirational Content: The speech contains powerful, motivational messages about personal empowerment and positive influence on others.


  • Concise Delivery: At 104 words, the speech is relatively brief, which can help maintain audience attention and focus on core ideas.






Areas for Improvement




  • Speaking Speed: At 130 words per minute (wpm), the pace is slightly below the ideal range of 140-200 wpm. Increasing the speed slightly could make the delivery more engaging.


  • Filler Words: The use of "ah" was noted in the transcript. Reducing or eliminating filler words can enhance clarity and professionalism.


  • Structure: While the content is inspirational, a clearer structure with an introduction, main points, and conclusion could improve overall coherence.






Summary Feedback



The speaker delivers a powerful message with good use of pauses, allowing the audience to absorb the inspirational content. To enhance the speech:




  1. Increase speaking speed slightly to reach the ideal range of 140-200 wpm.

  2. Eliminate filler words like "ah" to maintain a smooth flow of ideas.

  3. Consider structuring the speech more clearly, perhaps by grouping ideas into distinct sections.

  4. The personal touch at the end ("Sir, I just want to say thank you. You saved my life.") is impactful but seems abrupt. Consider a smoother transition or integration of this personal element.



By implementing these suggestions, the speaker can elevate an already powerful message to create an even more impactful and polished presentation.








Screenshot: Show's Transcript









The variours steps in the process Speech->Transript->Inference.




CODE
## All lines except those with leading ##'s are from server logs
## Detection of directive from first line
f(getComnand) from first FinalTranscript: I would like you to generate a catchy summary and sentiment analysis of the following story. {
summarization: true,
summary_model: 'catchy',
summary_type: 'gist',
sentiment_analysis: true
}

## The streamed audio packets are then combined with .wav header and
## file object for the wave file is uploaded

url : https://api.assemblyai.com/v2/upload

## we then have a URL to the uploaded audio
f(fileUpload) : Uploaded Successfully {
upload_url: 'https://cdn.assemblyai.com/upload/ebabc86a-0a1d-4988-a6f0-5dc06352921d'
}
f(fileUpload): 472.023ms
/api/upload fileUpload: 540.219ms

## NOTE: parameters for audioURL -> transcript
## disfluencies - so assemblyai can transcribing filler words
## Parameters below are computed for every speech and depend
## on first line - see below
## audio_start_from - ensure we don't include first line while asking
/api/upload getTranscriptFromURL: 6.628s

f(getTranscriptFromURL) {
audio: 'https://cdn.assemblyai.com/upload/ebabc86a-0a1d-4988-a6f0-5dc06352921d',
disfluencies: true,
summarization: true,
summary_model: 'catchy',
summary_type: 'gist',
sentiment_analysis: true,
audio_start_from: 10010
}

## Transcript id along with a CUSTOM PROMPT is submitted to leMUR for
## specfic speech evaluation response.
## CUSTOM PROMPT contains:
## 1. Description of the Task evaluation details,
## 2. Required outline of the evaluation report
## 3. Acceptable standards of a good speech
## 4. App generated metrics (duration, words per min, pauses, etc)
## 5. Report format as markdown

/api/feedback 1129a02c-4a28-49ea-8fbf-cb55f89a0214
f(askLeMUR): 9.234s










Screenshot: Show's Transcription and Sentiment Analysis




and @3kb-dev’s WAV file guide were lifesavers.
<!-- Please clearly note whether or not this submission should qualify for additional prompts, and if so, expand on how you implemented those additional tools. -->



  • Prompt Coverage.
    In my opinion SpeechTrack qualifies for all 3 challenge prompts as it ended up using:


  • Streaming Speech-to-Text - to provide realtime cues (providing the audio and real-time transcript)


  • Speech-to-Text - Transcribe Audio (with in-audio parameter detection)


  • Speech Understanding (using transcript of (b) and LeMUR to provide analysis and feedback on Speech)






  • Reflections & Conclusion



    Though the journey had its hurdles, completing SpeechTrack was immensely satisfying. Thanks to Assembly AI for their fantastic APIs, great support (thanks to Lee Vaughn of Support Engineering at Assembly AI) and the Dev community for their support. Here’s to more learning and creating in the future!

    Vollständiger Original-Bericht
    Ausführliche Details, Code-Beispiele & Hersteller-Stellungnahme auf dev.to.
    ↗ Original-Artikel auf dev.to lesen
    Wie bewertest du diesen Beitrag?
    1 Klick Feedback
    Teilen mit Netzwerk & Team:

    Community-Analysen & Experten-Meinungen 0

    Verfasse deine eigene Analyse, teile Workarounds oder diskutiere diesen Vorfall im Blog.
    Noch keine Community-Analyse verfasst. Markiere einen Textabschnitt oder klicke oben auf Eigene Analyse verfassen“!
    Community Pulse: Relevanz-Einschätzung
    1 Klick Experten-Votum
    🔴 Akute Relevanz 0%
    🟡 In Evaluierung 0%
    🟢 Keine Auswirkung 0%
    Spannende Innovation 0%
    Verwandte Story-Cluster & Quellen (Vektor-KI)
    Port 8095 Engine
    1 Quelle
    The Gemini desktop app is now available for Windows
    1 Quelle
    Burn Out, Or Fade Away
    1 Quelle
    Windows 11 KB5129195 is out after Microsoft confirms major issues with the September 2026 update, but it won’t fix AMD GPU errors
    Ähnliche Beiträge
    🔍 Verwandte News

    Auch interessante Nachrichten SpeechTrack

    Thematisch verwandte Begriffe: SpeechTrack · 6 Treffer

    Laden...

    Videos werden geladen ...

    Laden...

    Beiträge werden geladen ...

    Laden...

    Videos werden geladen ...