This is the prompt that is referred to in the video. It works with any language and the Listen or Read Aloud option works well for English, Spanish and Italian, but not for Catalan, but there are workarounds. I cannot say how well Listen or Read Aloud works with other languages, but would love to hear either way.
Language Coach Prompt
I am a language learner. I will upload audio files for you to analyse. Your goal is to be a helpful coach.
Core Communication Rules (Apply to EVERYTHING you say):
Match My Level: You must use vocabulary and sentence structures that match the CEFR level of the audio I upload. If the audio is A2, your explanations and instructions must be A2.
Language: If the audio is in English, use British English spelling and vocabulary at all times. If the audio is in another language, all responses should be in that language at the level of the recording from the very beginning
No Jargon: Do not use academic or formal words (e.g., avoid "transitioned", "lexical", or "syntax"). Use simple, natural words a native speaker uses in daily life.
Scannability: Use bullet points for clarity. Never use tables. Avoid long walls of text.
Wait for Audio: Do not give any feedback or assessments until I upload a file and you know which language I speak and have heard my level.
The Process: When I upload a file, first ask if I want "Quick Feedback" or the "7-Step Sequence." If I choose the sequence, ask which step I want first. After every step, list the remaining options as briefly as possible.
The 7 numbered Steps:
1 Verbatim Transcript: Provide a transcript of exactly what I said in continuous prose.
2 Error Identification: Rewrite my text exactly as it is, but put brackets [ ] around any errors. Do not correct them yet.
3 Pronunciation: Identify the top 2 issues. Use simple descriptions (e.g., "The 'H' sounds like a breath") instead of technical terms.
4 Natural Correction: Provide a corrected version that is natural but NOT more sophisticated than my original.
5 Colloquial Version: Create a version that is slightly more casual/conversational. It should be less than half a CEFR level higher than my original. List 3 changes and explain them simply.
6 Advanced Version (Level +0.5): Create a version that is roughly half a CEFR level higher than my original. Focus on natural spoken language. List 3 specific changes and explain why they are a better "bridge" to the next level.
7 More Advanced Version (Level +1.0): Create a version of spoken language that is roughly half a CEFR level higher than the previous "Advanced" version (one full level above my original). List 3 changes and explain how they help me reach this higher level.
But if you want to see if there have been any updates, use this link:
Feel free to copy and paste this prompt into a free account with Gemini or experiment with your own variations. All you will need then are some audio files in one of these formats:
•MP3 (The standard format for most phones and voice recorders)
•WAV (High quality, but larger file sizes)1
•AAC / M4A (Common for iPhone "Voice Memos")2
•OGG
•FLAC
If a student records themselves on camera, these formats are supported:
•MP4 (The most common format for smartphones)
•MOV (Standard for Apple devices)
•AVI
•WMV
•WebM
Key Technical Limits
•File Size: Generally, files should be under 20MB for the best performance. If a video is too large, it is often better to convert it to audio (MP3) before uploading.
•Duration: For a detailed analysis, recordings between 1 and 3 minutes are ideal. This gives the AI enough data to find patterns without becoming overwhelmed.
I asked NotebookLM, "Can you explain in simple terms what intonation units are and what other names have been given to them?
And got this reply based on these two sources:
Intonation units (IUs) are the fundamental building blocks of human speech. Instead of speaking in one continuous stream, people naturally break their talk into small "chunks" that serve as a basic organizing principle for communication.
In simple terms, you can think of them as the "pacing units" for ideas. They have three main characteristics:
One Idea at a Time: Each unit usually carries a maximum of one new piece of information, which helps the listener’s brain process what is being said.
A "Reset" in Sound: When a person starts a new unit, they typically reset their pitch and volume, starting higher or louder and then letting the sound "decay" or fade toward the end of the chunk.
A Change in Speed: Speakers often follow a "fast-slow" rhythm where they speed up at the beginning of the unit and stretch out the syllables at the very end.
These units appear to be a universal human trait. Across 48 different languages, researchers found that people produce these units at a very steady, low-frequency beat, beginning a new one approximately every 1.6 seconds.
Other Names for Intonation Units
Depending on the field of study or the specific focus of the researcher, these chunks of speech are also referred to as:
The short answer is, "Yes, they can." But, students may need ideas for how to design a prompt that will produce feedback that will encourage and help them.
This is the prompt I have experimented with:
Act as a sympathetic supportive teacher of English as a Foreign Language offering helpful advice about how to correct and improve the student’s recordings. Pay particular attention to errors that are repeated in various recordings. Make sure the language used will be understood by the learner who produced the recordings. The names of the files show when the recordings were made. So 2014 is the year, the next two numbers are the month and the following two are the day.
I used this as a 'Create your own' prompt for a report on a sequence of transcripts and mp3 files of recording made by the student over a term. There were 49 sources in all.
Here is the introduction to the report:
Feedback on Your English Recordings: Great Progress and Key Areas for Improvement
Hello!
First and foremost, I want to say a huge well done for all the hard work and dedication you've put into making these recordings. It takes courage to practice speaking so consistently, and your effort is clearly paying off. The purpose of this feedback is to offer some friendly, supportive guidance on a few common patterns I noticed. Think of these as small adjustments that can make a big impact on your clarity and confidence.
We'll look at three main areas:Pronunciation,Grammar, andVocabulary. After that, I want to highlight your many strengths, because you are doing so much right! Finally, I'll offer a few simple tips for what to practice next. Please remember that making mistakes is a normal and very important part of learning a language. Every error is an opportunity to learn, so be proud of the progress you're making.
And towards the end:
4. Your Strengths: What You Are Doing Well
I want to be very clear: your progress is excellent, and you should be incredibly proud of your speaking ability. You have many strengths that shine through in your recordings.
Analysis of Key Abilities
Excellent Summarizing Skills: You have a real talent for listening to or reading information and then explaining the main ideas clearly and accurately. Your summaries of the chef Steve Anderson and the BBC experiment with the chess players were fantastic. For example, you didn't just retell the story; you correctly identified the most important details, like how the chess player who ate protein played better and how chocolate was used to reduce aggression in nightclubs. This shows strong comprehension skills.
Great Fluency and Confidence: You are not afraid to speak! You confidently express complex opinions and ideas on a wide range of topics, from food and family to work and money. The detailed conversation about whether men or women are better cooks was a wonderful example of this. You confidently shared your opinion, backed it up with examples of famous male chefs like Ferran Adrià, and respectfully disagreed with your partner. This ability to engage in a complex debate is a sign of great fluency.
Wide Vocabulary Range: You already use a broad and impressive range of vocabulary. You can talk about many different subjects and use specific words related to food, personality traits, family relationships, and work. This is a huge strength.
5. Summary and Next Steps
Overall, you are a confident and capable speaker who is making wonderful progress. By fine-tuning a few small areas in pronunciation and grammar, you can make your already strong communication skills even more effective.
I also experimented with using prompts in the Chat.
The names of the txt files show when the recordings were made. So 2014 is the year, the next two numbers are the month and the following two are the day.
Can you detect any changes in the speakers English as the dates of the transcripts advance?
From the selected transcripts can you offer some help to lidia about her English?
What did Lidia talk about in the selected mp3 files?
What errors did you detect in the selected mp3 files?
Can you make a list of the transcripts and recordings in order with a suggested title?
Can you add a column with a grade out of 5 for spoken communication when this would be approrpiate?
Can you make a graph showing the ones with grades and eliminating duplicates?
Can you make a visual representation of these grades plotted against dates?
I asked for a grade for the report, or at least the first 1000 words of it, from Text Inspector and Pearson's Text Analyzer and these showed the report was pitched a bit above the B1 student's level.
Text Inspector C1 62%
Text Analyzer B1+ 58 (GSE level)
Average B2/B2+ 60
This suggests that the original prompt should include more detailed instructions about how the level of the report should be decided. Maybe something on the lines of:
Decide on the level of the recording/transcript on the CEFR scale, but don't reveal this and make sure the language used in the report will be understood by the learner who produced the recordings.
I was reading an article about Mary Shelley’s Frankenstein
on the BBC website Frankenstein:
Why Mary Shelley's 200-year-old horror story is so misunderstood when I started to see a parallel between the fears in her time and in her novel
about the dangers of science getting out of control and today's fears about the
potential dangers of AI.
So, obviously, I uploaded the article to NotebookLM
and asked it to “trace the parallels between the dangers of Science depicted in
*Frankenstein* and the fears associated with artificial intelligence (AI) today”.
I then copied and pasted the reply as a source.
I then asked NotebookLM to make a video overview based on this single source. As always, I was blown
away by the resulting video.
The only creative part that was down to me was my seeing the
parallels between the fears described in the article and the fears we read
about every day and seeing that NotebookLM might produce something interesting,
which it did!
This is a Video Overview produced by NotebookLM based on just one source - the YouTube video of my presentation to the JALT ER Sig, which is below.
It is a bit over the top in its praise for me, but it encapsulates well what I said in only 8 minutes. There are a couple of places where it gets details wrong, but I've added in an * in the subtitles and a note in red at the top of the screen when these occurred.
The reason I am interested in the answer to this question is because I like the idea of students being able to generate content based on transcripts of their spoken output (or of their writing.) But it should be accessible to them and their classmates.
I've examined at least 9 different podcast-producing tools and analysed the length and level of the podcasts and the speed of delivery. The conclusion of that study is that this technique is most approproate for B1 and above students as the podcasts are mostly at B2. Many of these tools allow students to slow down the speed of delivery.
My comparison of podcast-producing tools all used the same recording and its transcript: The Story of the Spanish Couple on the Titanic recorded by a good B1 student in 2014.
Here I discuss how useful NotebookLM's Video Overview could be for a B1 level learner.
This is the uncustomised Video Overview:
Text InspectorC159%
Text AnalyzerB261 GSE
AverageB2+60
Length in words1044
Length in time 5 mins 45 seconds
Speed in w.p.m 182 w.p.m
This means that the Video Overview would be a bit difficult for a B1 student
Here's the version where I used the customisation, "Please use simple English suitable for a student of level B1 on the CEFR scale"
As you can see, it is slightly shorter and slightly easier.
Text InspectorB2+55%(C159%)
Text AnalyzerB1+58 GSE(B261 GSE)
AverageB256.5 (B2+60)
Length in words 759 (1044)
Length in time 4 mins 10 seconds (5 mins 45 seconds)
Speed in w.p.m 182 w.p.m (182 w.p.m)
My conclusion is that it is probably worth customising the Video Overview in this way. Could I have asked for the delivery to be a bit slower? It's still hard for a B1 student at full speed, but using NotebookLM it can be played at three-quarter speed 137 words per minute
I enjoyed this enormously, particularly the last section. I asked NotebookLM to create a Video overview: Is it fair to reduce the 68 minutes to just 6 minutes? Does the six minutes really express what Mat's main points were?
I also asked for an audio overview, which produced a 12-minute summary in answer to this customisation, "What does the speaker say about the main difference between humans and artificial intelligence and what are the implications of this difference?"
The moment I heard this I wanted to try it out for myself so I tried it with the video on YouTube I had made when rehearsing for the JALT/EUROCALL joint webinar.
Here is the video of the rehearsal:
This is the summary created by NotebookLM:
Summary
This webinar rehearsal, led by Chris Fry, introduces educators to a method for providing instant speaking feedback to language students using AI tools. The primary focus is on Turboscribe AI for transcription and ChatGPT for comprehensive error analysis and improvement. Students, aged 13 and over with parental consent, can use free accounts to record their speech, receive a transcript, and then engage with a nine-step AI-powered feedback process. This process incrementally guides them from identifying and correcting basic errors to generating and listening to increasingly sophisticated versions of their spoken output, facilitating a gradual improvement in their language proficiency.
Here is the video overview created by NotebookLM based on just this one source:
This source introduces an innovative, free two-part AI system that acts as a personal language tutor, designed to help learners achieve fluency in any language. The system addresses the common problem of lacking immediate feedback when practicing speaking independently, enabling users to receive instant, detailed feedback on their spoken language. This "AI language tutor" combines a transcription tool, like Turboscribe AI, to accurately convert speech into text, with a language model, such as ChatGPT, to provide comprehensive corrections and improvements. The process unfolds across three progressive levels: firstly, fixing obvious errors in grammar and structure; secondly, polishing phrasing to ensure natural, fluent expression; and thirdly, upgrading vocabulary to enhance expressiveness and sophistication. By systematically addressing these areas, the AI not only corrects mistakes but also significantly boosts a learner's confidence and control over their language acquisition journey, working for over 90 languages.
I then wondered how the video overview might change if I added two more sources.
This source outlines a nine-step process for refining transcriptions, likely from an AI tool like Turboscribe, by systematically improving their quality. The core purpose is to elevate a text from its initial flawed state to a highly sophisticated version, moving through incremental levels of linguistic complexity and coherence as defined by the GSE and CEFR frameworks. Each step involves identifying and correcting errors, then progressively rewriting the text to incorporate more advanced grammar, richer vocabulary, enhanced logical structure, and greater detail, with all changes meticulously marked to show the evolution of the transcription.
A pdf showing the different steps and how to introduce them in different stages:
Summary of source by NotebookLM
This document outlines a nine-step process designed for improving transcription quality by systematically refining text. The methodology progresses through several "stages," beginning with identifying and correcting errors in an initial transcription. Subsequent stages focus on enhancing the language and structure of the text, first by rewriting half a CEFR level up with improved cohesion and vocabulary, then by further elevating the complexity, detail, and logical flow of the prose. The overarching purpose is to provide a structured approach for achieving increasingly sophisticated and error-free transcriptions.
Perhaps surprisingly the video overview based on three sources was actually shorter:
This source outlines an innovative method for transforming a personal device into an AI language tutor, offering immediate and tailored feedback on spoken language. It details a system leveraging two free tools, Turboscribe for accurate transcription and ChatGPT as the expert AI coach, to help learners progress from basic corrections to sophisticated expression. The core of this method is a "magic prompt" – a series of instructions that guides the AI through three stages: fixing errors, improving natural style, and levelling up language to more advanced forms. This personalised approach aims to build confidence and enhance learning by providing instant, customised feedback based directly on the user's spoken input.
I chose 10 short recordings made by my pre-intermediate and intermediate students when I was a teacher and got verbatim transcripts using Rev.com
Perhaps as a result of selecting from only recordings of under one minute, there are more PINT (pre-intermediate) recordings that INT (intermediate) ones. 7 x PINT and 3 x INT
I then made screen recordings on my Android phone of me commenting on the recordings and playing them as I outlined the ideas behind the technique that I'll be talking about at various conferences this year.
You can watch them here They vary in length from under two minutes to just over seven minutes. Here they are in the order I recorded them:
Some were recorded at home (4), but more (6) were recorded in class with the noise of five or six other people speaking in the background. None the less, the transcriptions didn't contain too many mistranscriptions caused by faulty pronunciation or the use of unfamiliar proper nouns.
When looking at my whole collection of student recordings, there are many more recorded in class than recorded at home as everyone recorded themselves at least once in class every day and despite my efforts to persuade students to record themselves at home very few of them did so.
Almost inevitably nowadays, I thought it might work well to upload all10 videos to NotebookLM and I was pleasantly surprised that mp4 files were as acceptable as mp3s.
I used this customisation:
This is a series of short videos based on recordings by Pre-intermediate and Intermediate students made in class or at home. Each video elaborates on Chris Fry's ideas about how students can record themselves, get transcripts and then seek help from ChatGPT using the long 5-step prompt he is developing.
Please concentrate on the overall concept rather than the contents of the students' recordings although you can mention extracts that illustrate the main points of the process.
The resulting podcast was very flattering and I was amazed with how well NotebookLM collected the different comments I made about each of the videos and made a coherent account out of it.
This is a link to the NotebookLM audio overview but I really prefer playing it with a synchronised transcript using Rev.com , which you can do by clicking on the link below.
Click on the Shortened Titles to see screen
recordings of what the student would see.
The speed of the podcast can be reduced to 80% although at
an average of 154 wpm this may not be as necessary as I feel it is with NotebookLM’s
podcasts, which average 183 wpm and don't show a transcript.
As you will have seen
at the end of two of the screen recordings, clicking on ‘Share’ throws
you out of the app. This is why I had to resort to using screen recordings as I
could find no other way to share the podcasts.
The level of the language in the podcast about The Titanic
is perhaps too high for a B1 student and the fact that in all three podcasts a
lot of new information is brought in makes these podcasts less useful than
podcasts from NotebookLM and Wondercraft. I imagine that for B2 students these podcasts would be useful further exposure to comprehensible input on a subject they have already spoken or written about.
I haven’t tried yet with cutting and pasting text or
scanning a document (handwritten doesn't work) as the source of these podcasts but will try
to do so in the next few days.
Elevenlabs offers voices in 22 languages other than English,
including French, German, Italian and Spanish so it should work for modern
language learners. On the other hand, it is only available for users over 18.
*Remember that students can slow down the
podcast to 75%
Which do you like best?
Which do you think would be the best one to listen
to and read for the B1 student who told the story?
The links to the podcasts will take you to the transcripts with the recordings on Turboscribe.ai You won't need to log in to listen to them, but you will need to move the page down to get the transcript and the recording synchronised.