Showing posts with label #Deepseek. Show all posts
Showing posts with label #Deepseek. Show all posts

Wednesday, 4 March 2026

Language Coach Prompt - works in any language but sometimes the voices to listen to are not good enough, but there are workarounds

The prompt works with any language and the Listen or Read  Aloud option works well for English, Spanish and Italian, but not for Catalan, but there are workarounds.

Here are some that I have tried with Catalan:

  1. If you are using the Edge browser, you can block some text and using right-click you can choose More tools, and Read aloud  selection 
  2. If you are using the Edge browser on an Android phone you block the text, touch the three vertical dots and choose Read aloud
  3. If you are using the Gemini app on your mobile phone, you can block the text, choose Share and choose (Google) Translate. Then touch the Speaker icon
  4. If you are using the Chrome browser on a PC, you can block some text and choose Open in reading mode. I had to choose CatalĂ  from the Voice selection icon
  5. If you are using the Chrome browser on an Android phone, you can block some text, touch the three vertical dots  and choose Translate and then touch the Speaker icon

Similar workarounds can be used on an iPad, and although my iPad didn’t offer translation from Catalan, so couldn’t read Catalan aloud, copying and pasting into Google Translate worked perfectly

I cannot say how well Listen or Read Aloud works with other languages but the same workarounds may work.  If you try it I would love to hear about how well it worked.

Trying Language Coach Prompt with different LLMs

Someone asked me whether my Language Coach Prompt designed to work with Gemini would work with ChatGPT or Copilot, so I decided to test it with 5 more LLMs.

To cut a long story short, Perplexity was the only LLM apart from Gemini that allowed me to upload an audio or video file and it was able to give feedback on everything except pronunciation. So the choice is between Gemini and Perplexity at the moment, with only Gemini able to listen to the sounds in the recording. See the end for details of this difference.

Trying it with ChatGPT

You’re out of uploads — try a Business plan for more. Otherwise, your limit will reset tomorrow after 7:19 PM.

Trying it with Copilot

This file type is not supported. See: File formats supported by Copilot

Upload a different file. No audio or video file formats are allowed. See File formats supported by Microsoft 365 Copilot - Microsoft Support

Trying it with Claude

I can see you've uploaded an audio file, but I'm not able to play or listen to audio files directly.

Could you try one of these:

  • Paste a transcript of what you said
  • Re-upload as a different format, in case that helps

I'm ready to help as soon as I can read or hear your words!

Trying it with Deepseek

I asked Deepseek, "Can I upload an audio or video file?" and got the reply:

Yes, you can upload an audio file directly to the chat.

Please use the upload button to attach your file. Once you have done that, we can begin.

(Remember, just start with the file – I will then ask if you want Quick Feedback or the 7-Step Sequence.)

I then said, "It says text extraction only. There's no mention of audio files." and got the reply:

You are absolutely right, and I apologise for the confusion. Thank you for checking that.

My ability to "hear" audio depends on the platform. Here, I can only read text that is typed or uploaded as a document.

Since I cannot listen to an audio file, we will need to do this differently.

Here are the two best ways to move forward:

    1. Type your speech: You can write out what you would have said in the audio. For example, write a paragraph as if you were speaking to a friend.
    2. Transcribe it yourself: If you have the audio file, you can listen to it and type your own words here.

Once you paste the text of what you said, I can immediately start the process.

Again, my apologies for the misunderstanding. Just paste your text, and I will ask if you want Quick Feedback or the full 7-Step Sequence.

Trying it with Perplexity

Do you want Quick Feedback or the 7-Step Sequence for this recording?

If you choose the 7-Step Sequence, which step should we do first?

  1. Verbatim Transcript
  2. Error Identification
  3. Pronunciation
  4. Natural Correction
  5. Colloquial Version
  6. Advanced Version (+0.5)
  7. More Advanced Version (+1.0)

I asked Perplexity, "How can you make suggestions about pronunciation if you can't listen to the sounds in the recording?" and got the reply:

You’re right to question that. I need to be clear and honest with you.

Right now, I cannot actually hear your audio. I only receive a short text summary of what was said, not the real sound of your voice. That means:

    • cannot truly judge your pronunciation.
    • cannot hear your accent, intonation, stress, or individual sounds.

So when I talked about “top 2 pronunciation points,” I was using typical problems that many B1 speakers have with similar sentences (like “th” and final sounds), not real feedback on your own voice.

Comparing Gemini’s Reply with Perplexity’s confession

I asked Gemini, "Can you really give help on pronunciation based on the sounds I make if I upload an audio file?" and got the reply:

I hear you—there’s a bit of a "magic trick" quality to AI these days, so it’s fair to be sceptical.

The short answer is yes, but with a slight distinction in how we do it. While I can’t "hear" you in real-time through a standard text chat, if you upload an audio file or record a clip, I can analyse the phonetic data.


How I Can Help

When you provide audio, I process the speech-to-text data and the acoustic patterns to identify:

    • Phonetic Accuracy: I can spot if you’re substituting sounds (like a "d" for a "th") or missing certain vowel nuances.
    • Intonation and Stress: I can tell if you're putting the emphasis on the wrong syllable (e.g., saying PHOtograph vs. phoTOGrapher).
    • Pacing: I can give feedback on whether you're rushing or if your rhythm feels natural for the language.

….

Sunday, 29 December 2024

An experiment with a 5-step prompt with 5 types of GenAI

I decided that it must be possible to include all the prompts for a routine sequence of prompts in one rather long prompt.

I wanted to get:

  1. A student transcript with the errors marked so they could try to see their errors
  2. A corrected version of the transcript
  3. An improved version at the students level (A2, in this case)
  4. An improved version at the next level (B1, in this case)
  5. An improved version at two levels up (B2, in this case)

The secret was to include the five steps in the prompt but to instruct GenAI to wait for the prompt “Next” before moving on to the next step.

I then took a short recording made by a very good A2 student and used Rev.com to get the transcript. The idea is that the student has to copy the long prompt from a WhatsApp group, for example, and paste it into the chosen GenAI and then copy the transcript from Turboscribe.ai or Rev.com  and paste it into the same GenAI.

I did this with these 5 types of GenAI:

  1. ChatGPT
  2. Gemini
  3. Claude
  4. Copilot
  5. Deepseek

I then copied the output from each GenAI for each step and pasted it into Pearson’s GSE Text Analyzer ( https://www.english.com/gse/teacher-toolkit/user/textanalyzer ), which gave me GSE and CEFR levels for all 25 versions of the transcript.

Apart from looking carefully at the English used in each version to try to understand what language had determined the levels given for them all, I also made an Excel spreadsheet with the two sets of levels. With these I was able to produce two graphs. The first one shows what happened with each type of GenAI:


At first glance Claude was the best at producing increasingly sophisticated versions of the student’s transcript, although they were always half a CEFR level too high. Mark you, the original transcript was already half a level higher as she was a very good student.

Equally obvious is the fact that Copilot was useless!

ChatGPT, Gemini and Deepseek failed to produce increasingly sophisticated versions across the four levels, so they didn’t do what I had intended.

In fact, as can be seen in this second chart, the whole idea didn’t work on average:


Once again, this may be as a result of the student’s original transcript being higher than expected (A2+ rather than A2). Maybe the lesson to be learnt from this is that instead of using fixed levels for each step up, the different GenAIs may be able to produce new version of the transcript one level higher on the CEFR scale. This would have the added advantage that the same long prompt could be useful for all students in a class and across all courses.

Here is the original version of the long prompt derived with help from Turboscribe and ChatGPT’s advice. (I hope the concept of including steps and a trigger word will be useful):

“You will be provided a transcript as well as instructions about what to do with that transcript. Unless otherwise specified, your response should be in the same language as the transcript.[1] Here are your instructions, which you must follow:

Instructions:

  1. I will provide a transcript. For each step, you should follow the instructions carefully.
  2. For each task, do not include timestamps unless specifically requested.
  3. After each step, wait for me to say "next" before proceeding. 

Steps:

  1. Mark the errors in the transcript in bold, but do not correct them.
  2. Correct the errors, marking the changes in bold and leaving the original errors in brackets.
  3. Improve the transcript for A2 level students, marking improvements in bold.
  4. Improve the transcript to a B1 level, marking changes in bold.
  5. Finally, enhance the transcript to a B2 level, marking improvements in bold. 

[Sample transcript of an A2 student’s recording:]

The history start in the summer of 2011. Anna went on holiday with some friends on island. The photo was taken on hill. Hill is a little mountain and called Ana. The photo is important for her because the stone is a mysterious for her. And she put your hands around it, the stone, and, and she was sleeping. And this photo there are in other places because in your mobile phone, computer, et cetera.”


[1] These two sentences came from Turboscribe’s Custom prompt