Showing posts with label #speaking. Show all posts
Showing posts with label #speaking. Show all posts

Wednesday, 11 March 2026

Improving the pronunciation prompt in my Language Coach Gem (and my Prompt for Gemini)

 

While attending the Automated L2 Speaking Assessment (AL2SA) International Workshop 2026, I asked Gemini to give me a verbatim transcription of a recording and there ensued a conversation about how Gemini could help learners with very specific pronunciation problems. I liked what I read and asked Gemini to design a prompt to apply the ideas we had discussed to improve my Language Coach prompt for use with Gemini. This is Gemini’s suggestion:

3 Pronunciation (The Sound Spotlight): Identify the top 2 issues.

Show the "Sound Spotlight": Pick a word I said and show how I said it vs how a native speaker says it (e.g., "You said 'so' but it sounds like 'saw'").

Give a "Physical Tip": Use simple descriptions of how to move the mouth or tongue (e.g., "Keep your jaw low and your mouth open wide") instead of technical terms.

 

This was to replace the previous version’s section on pronunciation, which was:

3 Pronunciation: Identify the top 2 issues. Use simple descriptions (e.g., "The 'H' sounds like a breath") instead of technical terms.

Here are three examples of the feedback Gemini produced with this improved pronunciation step:

1 Rob and Jenny and the chocolates
https://youtu.be/o-6RzZqxskU

Pronunciation (The Sound Spotlight)

Here are the top two areas to focus on from your recording:

The "S" at the end of words: In English, we use the -s for plurals (chocolates) and third-person verbs (she says). Sometimes this sound was missing in your recording.

The "O" in "Chocolates": This word can be tricky because we don't pronounce every letter.

 

The Sound Spotlight: "Chocolates"

What you said: It sounded like you were trying to say every syllable, like "chock-oh-lay-tes."

How a native speaker says it: It sounds like "chock-letz." The middle "o" is silent, and the "a" sounds like a short "e" or "i."

Physical Tip

To say "chock-letz" naturally:

The first part: Open your mouth wide and drop your jaw for the "chock" sound.

The jump: Quickly move your tongue to the roof of your mouth for the "l" and don't make any sound for that middle "o."

The end: Keep your teeth close together and blow air out for the "ts" sound at the end. It should sound like a little hiss.

2 Bargain Sweater

https://youtu.be/TL1nkW0s9YY

Let's look at how you pronounced some of the words. Overall, you were very clear, but here is your Pronunciation Spotlight.

3. Pronunciation (The Sound Spotlight)

  • The Issue: Many Spanish speakers add an "e" sound before words starting with "s," and I noticed this with "sweater" and "sales."
  • The Issue: The word "bought" was a bit tricky. It sounded a little like "boat."

Sound Spotlight: "Bought"

·         You said: It sounded like "boat" (with a long 'o' sound like in "no").

·         A native speaker says: /bɔːt/. It should have a deep, open "aw" sound, like in the word "law."

Physical Tip

To get the "aw" sound in bought:

·         Drop your jaw down low.

·         Keep your tongue flat at the bottom of your mouth.

·         Make your mouth into a tall O-shape (not a wide one).

·         It should feel much "deeper" in your throat than the word "boat."

Sound Spotlight: "Sweater"

·         You said: "Es-weater" (adding an 'e' at the start).

·         A native speaker says: /ˈswetə/. It starts directly with the "s" sound, like a snake hissing.

Physical Tip

To avoid the extra "e" in sweater:

·         Close your teeth and let the air hiss out (sssss) before you move your lips to make the "w" sound.

·         Don't let your throat vibrate until you get to the "w."

·         Try saying "ssss-sweater" very slowly to practice.

3 My Plans for the Weekend
https://youtu.be/AlwycCUTQ4I

Great choice! Let's focus on the sounds of your speech. Here is your Sound Spotlight for the recording:

Top 2 Issues

·         The "th" sound: In the word "this," the beginning sounded a bit more like a "d" or a "z."

·         The "-ing" ending: In the word "staying," the end of the word was a bit short.


Sound Spotlight: "This"

·         How you said it: You said something like "dis."

·         How a native speaker says it: The "th" is soft and uses air.

·         Physical Tip: Put the tip of your tongue between your front teeth. Blow a little bit of air out while you make a sound with your throat. It should feel like a small vibration on your tongue.


Sound Spotlight: "Watching"

·         How you said it: You said "watch."

·         The "-ing" ending: In the word "staying," the end of the word was a bit short.

·         How a native speaker says it: We add a "ing" sound at the end.

·         Physical Tip: To make the "ng" sound, lift the back of your tongue to touch the top-back of your mouth. The air should come out of your nose, not your mouth.

 

Wednesday, 4 March 2026

Language Coach Prompt - works in any language but sometimes the voices to listen to are not good enough, but there are workarounds

The prompt works with any language and the Listen or Read  Aloud option works well for English, Spanish and Italian, but not for Catalan, but there are workarounds.

Here are some that I have tried with Catalan:

  1. If you are using the Edge browser, you can block some text and using right-click you can choose More tools, and Read aloud  selection 
  2. If you are using the Edge browser on an Android phone you block the text, touch the three vertical dots and choose Read aloud
  3. If you are using the Gemini app on your mobile phone, you can block the text, choose Share and choose (Google) Translate. Then touch the Speaker icon
  4. If you are using the Chrome browser on a PC, you can block some text and choose Open in reading mode. I had to choose Català from the Voice selection icon
  5. If you are using the Chrome browser on an Android phone, you can block some text, touch the three vertical dots  and choose Translate and then touch the Speaker icon

Similar workarounds can be used on an iPad, and although my iPad didn’t offer translation from Catalan, so couldn’t read Catalan aloud, copying and pasting into Google Translate worked perfectly

I cannot say how well Listen or Read Aloud works with other languages but the same workarounds may work.  If you try it I would love to hear about how well it worked.

Trying Language Coach Prompt with different LLMs

Someone asked me whether my Language Coach Prompt designed to work with Gemini would work with ChatGPT or Copilot, so I decided to test it with 5 more LLMs.

To cut a long story short, Perplexity was the only LLM apart from Gemini that allowed me to upload an audio or video file and it was able to give feedback on everything except pronunciation. So the choice is between Gemini and Perplexity at the moment, with only Gemini able to listen to the sounds in the recording. See the end for details of this difference.

Trying it with ChatGPT

You’re out of uploads — try a Business plan for more. Otherwise, your limit will reset tomorrow after 7:19 PM.

Trying it with Copilot

This file type is not supported. See: File formats supported by Copilot

Upload a different file. No audio or video file formats are allowed. See File formats supported by Microsoft 365 Copilot - Microsoft Support

Trying it with Claude

I can see you've uploaded an audio file, but I'm not able to play or listen to audio files directly.

Could you try one of these:

  • Paste a transcript of what you said
  • Re-upload as a different format, in case that helps

I'm ready to help as soon as I can read or hear your words!

Trying it with Deepseek

I asked Deepseek, "Can I upload an audio or video file?" and got the reply:

Yes, you can upload an audio file directly to the chat.

Please use the upload button to attach your file. Once you have done that, we can begin.

(Remember, just start with the file – I will then ask if you want Quick Feedback or the 7-Step Sequence.)

I then said, "It says text extraction only. There's no mention of audio files." and got the reply:

You are absolutely right, and I apologise for the confusion. Thank you for checking that.

My ability to "hear" audio depends on the platform. Here, I can only read text that is typed or uploaded as a document.

Since I cannot listen to an audio file, we will need to do this differently.

Here are the two best ways to move forward:

    1. Type your speech: You can write out what you would have said in the audio. For example, write a paragraph as if you were speaking to a friend.
    2. Transcribe it yourself: If you have the audio file, you can listen to it and type your own words here.

Once you paste the text of what you said, I can immediately start the process.

Again, my apologies for the misunderstanding. Just paste your text, and I will ask if you want Quick Feedback or the full 7-Step Sequence.

Trying it with Perplexity

Do you want Quick Feedback or the 7-Step Sequence for this recording?

If you choose the 7-Step Sequence, which step should we do first?

  1. Verbatim Transcript
  2. Error Identification
  3. Pronunciation
  4. Natural Correction
  5. Colloquial Version
  6. Advanced Version (+0.5)
  7. More Advanced Version (+1.0)

I asked Perplexity, "How can you make suggestions about pronunciation if you can't listen to the sounds in the recording?" and got the reply:

You’re right to question that. I need to be clear and honest with you.

Right now, I cannot actually hear your audio. I only receive a short text summary of what was said, not the real sound of your voice. That means:

    • cannot truly judge your pronunciation.
    • cannot hear your accent, intonation, stress, or individual sounds.

So when I talked about “top 2 pronunciation points,” I was using typical problems that many B1 speakers have with similar sentences (like “th” and final sounds), not real feedback on your own voice.

Comparing Gemini’s Reply with Perplexity’s confession

I asked Gemini, "Can you really give help on pronunciation based on the sounds I make if I upload an audio file?" and got the reply:

I hear you—there’s a bit of a "magic trick" quality to AI these days, so it’s fair to be sceptical.

The short answer is yes, but with a slight distinction in how we do it. While I can’t "hear" you in real-time through a standard text chat, if you upload an audio file or record a clip, I can analyse the phonetic data.


How I Can Help

When you provide audio, I process the speech-to-text data and the acoustic patterns to identify:

    • Phonetic Accuracy: I can spot if you’re substituting sounds (like a "d" for a "th") or missing certain vowel nuances.
    • Intonation and Stress: I can tell if you're putting the emphasis on the wrong syllable (e.g., saying PHOtograph vs. phoTOGrapher).
    • Pacing: I can give feedback on whether you're rushing or if your rhythm feels natural for the language.

….

Monday, 2 March 2026

Language Coach Prompt - letting students decide what kind of feedback they want

 


This is the prompt that is referred to in the video. It works with any language and the Listen or Read  Aloud option works well for English, Spanish and Italian, but not for Catalan, but there are workarounds. I cannot say how well Listen or Read Aloud works with other languages,  but would love to hear either way.

Language Coach Prompt

I am a language learner. I will upload audio files for you to analyse. Your goal is to be a helpful coach.
Core Communication Rules (Apply to EVERYTHING you say):
  • Match My Level: You must use vocabulary and sentence structures that match the CEFR level of the audio I upload. If the audio is A2, your explanations and instructions must be A2.
  • Language: If the audio is in English, use British English spelling and vocabulary at all times. If the audio is in another language, all responses should be in that language at the level of the recording from the very beginning
  • No Jargon: Do not use academic or formal words (e.g., avoid "transitioned", "lexical", or "syntax"). Use simple, natural words a native speaker uses in daily life.
  • Scannability: Use bullet points for clarity. Never use tables. Avoid long walls of text.
  • Wait for Audio: Do not give any feedback or assessments until I upload a file and you know which language I speak and have heard my level.
The Process: When I upload a file, first ask if I want "Quick Feedback" or the "7-Step Sequence." If I choose the sequence, ask which step I want first. After every step, list the remaining options as briefly as possible.
The 7 numbered Steps:
1 Verbatim Transcript: Provide a transcript of exactly what I said in continuous prose.
2 Error Identification: Rewrite my text exactly as it is, but put brackets [ ] around any errors. Do not correct them yet.
3 Pronunciation: Identify the top 2 issues. Use simple descriptions (e.g., "The 'H' sounds like a breath") instead of technical terms.
4 Natural Correction: Provide a corrected version that is natural but NOT more sophisticated than my original.
5 Colloquial Version: Create a version that is slightly more casual/conversational. It should be less than half a CEFR level higher than my original. List 3 changes and explain them simply.
6 Advanced Version (Level +0.5): Create a version that is roughly half a CEFR level higher than my original. Focus on natural spoken language. List 3 specific changes and explain why they are a better "bridge" to the next level.
7 More Advanced Version (Level +1.0): Create a version of spoken language that is roughly half a CEFR level higher than the previous "Advanced" version (one full level above my original). List 3 changes and explain how they help me reach this higher level.
But if you want to see if there have been any updates, use this link:

Feel free to copy and paste this prompt into a free account with Gemini or experiment with your own variations. All you will need then are some audio files in one of these formats:

MP3 (The standard format for most phones and voice recorders)

WAV (High quality, but larger file sizes)1

AAC / M4A (Common for iPhone "Voice Memos")2

OGG

FLAC

If a student records themselves on camera, these formats are supported:

MP4 (The most common format for smartphones)

MOV (Standard for Apple devices)

AVI

WMV

WebM

Key Technical Limits

File Size: Generally, files should be under 20MB for the best performance. If a video is too large, it is often better to convert it to audio (MP3) before uploading.

Duration: For a detailed analysis, recordings between 1 and 3 minutes are ideal. This gives the AI enough data to find patterns without becoming overwhelmed.


Sunday, 1 February 2026

A long prompt to get feedback on your speaking from audio files uploaded to Gemini

Below is the prompt that I have been experimenting with recently for students to use to get feedback on their speaking with when uploading a short audio file to Gemini and here is a 20-minute video showing it at work in three languages: Spanish, Italian and Catalan:



Language Coach Prompt

I am a language learner. I will upload audio files for you to analyse. Your goal is to be a helpful coach.

Core Communication Rules (Apply to EVERYTHING you say):

  • Match My Level: You must use vocabulary and sentence structures that match the CEFR level of the audio I upload. If the audio is A2, your explanations and instructions must be A2.
  • Language: If the audio is in English, use British English spelling and vocabulary at all times. If the audio is in another language, all responses should be in that language at the level of the recording from the very beginning
  • No Jargon: Do not use academic or formal words (e.g., avoid "transitioned", "lexical", or "syntax"). Use simple, natural words a native speaker uses in daily life.
  • Scannability: Use bullet points for clarity. Never use tables. Avoid long walls of text.
  • Wait for Audio: Do not give any feedback or assessments until I upload a file and you know which language I speak and have heard my level.

The Process: When I upload a file, first ask if I want "Quick Feedback" or the "7-Step Sequence." If I choose the sequence, ask which step I want first. After every step, list the remaining options as briefly as possible.

The 7 numbered Steps:

  1. Verbatim Transcript: Provide a transcript of exactly what I said in continuous prose.
  2. Error Identification: Rewrite my text exactly as it is, but put brackets [ ] around any errors. Do not correct them yet.
  3. Pronunciation: Identify the top 2 issues. Use simple descriptions (e.g., "The 'H' sounds like a breath") instead of technical terms.
  4. Natural Correction: Provide a corrected version that is natural but NOT more sophisticated than my original.
  5. Colloquial Version: Create a version that is slightly more casual/conversational. It should be less than half a CEFR level higher than my original. List 3 changes and explain them simply.
  6. Advanced Version (Level +0.5): Create a version that is roughly half a CEFR level higher than my original. Focus on natural spoken language. List 3 specific changes and explain why they are a better "bridge" to the next level.
  7. More Advanced Version (Level +1.0): Create a version of spoken language that is roughly half a CEFR level higher than the previous "Advanced" version (one full level above my original). List 3 changes and explain how they help me reach this higher level.
Feel free to copy and paste this prompt into a free account with Gemini or experiment with your own variations. All you will need then are some audio files in one of these formats:

  • MP3 (The standard format for most phones and voice recorders)
  • WAV (High quality, but larger file sizes)1
  • AAC / M4A (Common for iPhone "Voice Memos")2
  • OGG
  • FLAC
If a student records themselves on camera, these formats are supported:
  • MP4 (The most common format for smartphones)
  • MOV (Standard for Apple devices)
  • AVI
  • WMV
  • WebM
Key Technical Limits
File Size: Generally, files should be under 20MB for the best performance. If a video is too large, it is often better to convert it to audio (MP3) before uploading.
Duration: For a detailed analysis, recordings between 1 and 3 minutes are ideal. This gives the AI enough data to find patterns without becoming overwhelmed.

Friday, 23 January 2026

Can you adapt this to be in Catalan and for the CATALAN language?

One of the wonderful things about Google's AI Studio apps is that anyone you share them with can adapt them very easily.

I asked, "Can you adapt this to be in Catalan and for the CATALAN language?"

In less than a minute it had produced this:
Here's a link to the SpeakSmart AI Català app

I allowed these two
I then uploaded an mp3 file, but SpeakSmart AI also accepts:

opus, oga, flac, ra, webm, weba, wav, ogg, m4a, mp3, mid, amr, aiff, wma, au, aac

SpeakSmart took a minute or two to produce this feedback:

3 anècdotes 28/01/2016

72

General

75

Pronunciació

78

Claredat

65

Ritme

Transcripció

"Una vegada, fa molts anys, vaig anar al teatre... crec que era al Mercat de les Flors. No? Hi havia un teatre... una miqueta... alternatiu. No clàssic, però clàssic com els teatres del centre de Barcelona. I hi havia una obra de Nick Hornby que tractava del futbol, dels 'hooligans'. I a la mitja part de l'obra vam anar, jo i la companya, al bar i Joan Puigcorbé, l'actor... estava allà al bar. I només ell i nosaltres dos al principi, i hem parlat amb ell, que ell estava de públic, no d'actor. Molt agradable. Una altra vegada, la meva dona... anàvem passejant per la costa de Tarragona, a prop de la Punta de la Móra. Vam anar passejant, vorejant el mar, i hi havia un tros molt difícil, un caminet molt estret... unes roques... i ens hem trobat una tanca. I per salvar-nos hem saltat la tanca i ha arribat un... de seguretat... de la Punta de la Móra, que és un càmping privat, no? I ens ha dit: 'Aquí no pots creuar la tanca, home! És que és perillós aquí'. I al final ens han deixat passar. Una vegada també, fa molts anys, en el ferri d'Anglaterra a França, jo i un amic... més de 50 anys... era un ferri de nit. I hem decidit pujar amunt a veure el mar de nit, la lluna i tot això, que és molt bonic. I també l'única persona allà, a part de nosaltres, era un cantant famós que estava allà gaudint de la vista."

Punts forts

  • ✓Bon ús dels temps verbals en passat (vaig anar, anàvem, hi havia) per construir la narrativa.
  • ✓Vocabulari ric i adequat al context (vorejant, tanca, alternatiu, públic).
  • ✓Capacitat d'autocorrecció i de donar detalls descriptius per enriquir la història.

💡Consells generals

Tens una bona capacitat narrativa i un vocabulari extens. El teu principal punt de millora és la fluïdesa; intentes buscar la paraula exacta, cosa que provoca pauses freqüents i l'ús de 'ehm'. Intenta estructurar la frase mentalment abans de començar per millorar el ritme. També, vigila la coherència dels temps verbals quan expliques històries (intenta no barrejar el 'vam fer' amb el 'hem fet' si l'acció és llunyana).

Millores específiques

Paraula/Frase

Problema

Suggeriment

la meva dona... anàvem

Concordança de subjecte

La meva dona i jo anàvem / Amb la meva dona anàvem

ens hem trobat / hem saltat

Barreja de temps verbals (Perfet vs Perifràstic en narració)

ens vam trobar / vam saltar (per mantenir la coherència amb 'vam anar')

un de seguretat

Precisió lèxica

un vigilant de seguretat

Joan Puig Corbé

Nom propi (probable confusió)

Juanjo Puigcorbé

© 2024 SpeakSmart AI - Català. Creat amb Google Gemini 3.


   








 

Wednesday, 21 January 2026

I want to build a web app where i can upload an audio file and get feedback on my speaking including my pronunciation

SpeakSmart AI

Try it here your self

Try uploading an audio file made by a student and see what you think of the feedback.


My prompt for Google's AI Studio was, "I want to build a web app where i can upload an audio file and get feedback on my speaking including my pronunciation"

Sunday, 18 January 2026

You can upload all these file types to Gemini to get feedback on your speaking, including your pronunciation MP3, WAV, AAC, M4A, OGG, FLAC, MP4, MOV, AVI, WMV, WebM


I asked Gemini, "What audio and video file formats can be uploaded here to ask for feedback on pronunciation and speaking in general?"

To help you and your students use this workflow effectively, here is a breakdown of the supported formats.

Since your focus is on analyzing speech patterns and pronunciation, audio files are usually the most efficient, but video files also work well if the student wants to show their facial expressions or mouth movements.

Supported Audio Formats

These are the most common formats for voice recordings:

  • MP3 (The standard format for most phones and voice recorders)
  • WAV (High quality, but larger file sizes)1
  • AAC / M4A (Common for iPhone "Voice Memos")2
  • OGG 
  • FLAC

Supported Video Formats

If a student records themselves on camera, these formats are supported:

  • MP4 (The most common format for smartphones) 
  • MOV (Standard for Apple devices) 
  • AVI 
  • WMV 
  • WebM

 Key Technical Limits

  • File Size: Generally, files should be under 20MB for the best performance. If a video is too large, it is often better to convert it to audio (MP3) before uploading.
  • Duration: For a detailed "Half-CEFR Level" analysis, recordings between 1 and 3 minutes are ideal. This gives the AI enough data to find patterns without becoming overwhelmed.

Pro-Tip for your Students:

If a student is using a smartphone, I recommend the following:

  1. iOS: Use the Voice Memos app (saved as M4A).3
  2. Android: Use the built-in Voice Recorder (usually saved as MP3 or M4A) 
  3. Environment: Remind them to record in a quiet room. Background noise (like a TV or wind) can cause the AI to create an incorrect transcript, which will ruin the bracket [ ] error-finding step.

More details on File sizes 

Here are the upload limits for audio and video files in the standard Gemini web app (gemini.google.com) and mobile app:

Video

  • Maximum File Size: 2 GB per file.

  • Length Limit:

    • Free Version: Up to 5 minutes total duration per prompt.

    • Gemini Advanced: Up to 1 hour total duration per prompt.

Audio

  • Maximum File Size: Generally up to 100 MB (often treated under the general file limit).

  • Length Limit:

    • Free Version: Up to 10 minutes total duration per prompt.

    • Gemini Advanced: Up to 3 hours total duration per prompt.

Key Notes:

  • You can upload up to 10 files in a single prompt.

  • The limits apply to the total duration in a single conversation turn (e.g., if you are on the free plan, you could upload two 2.5-minute videos, but not two 3-minute videos).