I. INTRODUCTION
From the perspective of utilizing movie English in language learning, it has been widely recognized that visual media such as movies and TV dramas can serve as highly effective materials for learning English. One major advantage is that movies include visual elements, which distinguishes them from written materials such as novels or newspapers. These visuals allow English as a Second Language (ESL) learners to observe speakers’ facial expressions and physical distance from one another. As a result, it becomes easier to understand the speaker’s emotional or mental state and also to grasp the reasons behind the use of demonstratives such as “this” (for nearby objects) and “that” (for distant ones) based on visual cues. Furthermore, since movies also provide spoken English, ESL learners can listen to conversations and acquire knowledge of actual English pronunciation in a natural context.
To date, numerous studies have explored the potential of movies and TV dramas as effective materials for English language education, along with reports on classroom practices. Specifically, there have been attempts to use dialogues for teaching English listening skills (
Hirano & Matsumoto, 2018;
Kondo, 2015), as well as efforts to help learners better understand English intonation (
Tabata, 2021). In addition, movie dialogues have been utilized to strengthen vocabulary in English and enhance the understanding of English grammar: for instance, through learning phrasal verbs (
Matsumoto, 2014) or teaching verb tense and aspect (
Kim & Kim, 2018). Moreover, some studies have incorporated these media into active learning strategies (
Kadoyama, 2019). What should be emphasized in these previous efforts is that, in all cases, the use of movies and TV dramas has been reported to benefit ESL learners and to yield positive educational outcomes. Indeed, many earlier studies have also reported that such use has a motivational effect on learners’ attitudes toward English learning (e.g.,
Hirose, 1997;
Kondo, 2018).
From the perspective of language education, efforts have been made to develop assessment tools that can efficiently and effortlessly measure learners’ language proficiency within a limited amount of time. For instance, based on the Simple Performance-Oriented Test (SPOT), which was designed to assess Japanese language proficiency in a short amount of time (
Kobayashi et al., 1995),
Maki et al. (2003) developed the Minimal English Test (MET), which enables rapid assessment of English proficiency. The MET is a test in which learners listen to around five minutes of spoken English and are required to fill in blanks in written sentences with appropriate English words consisting of no more than four letters.
Maki et al. (2003) found a significant correlation between MET scores and scores on the Center Test, a former university entrance examination in Japan (
r = .68).
1 Since then, the Maki Group have developed various versions of the MET, arguing that this test can provide a reasonably accurate prediction of performance on comprehensive English proficiency tests such as the Center Test and TOEIC, even within a very limited testing time.
However, there remains room for further development of the MET. For example, in the original versions of the MET, most of the English audio used features professional native-speaking narrators reading scripted sentences fluently, at a steady pace, and in the absence of any background noise. While it is common for English teaching materials and proficiency tests to employ carefully adjusted, easy-to-understand audio, continuing to rely solely on such idealized speech in English education raises important concerns. In reality, conversations rarely take place in such ideal conditions, nor do they always involve perfectly polished English.
Building on the discussion above, this study argues for the necessity of developing an English proficiency test that incorporates audio materials more reflective of real-world English usage, in order to enhance learners’ practical language skills and improve their communicative competence in English. In this regard, movie dialogues—which often include background noise, emotionally influenced speech (e.g., expressions of joy, anger, sadness), and a wide range of auditory challenges—can serve as promising material for such a test. Therefore, this study proposes the development of a movie-based version of the MET, referred to as the movie MET (mMET). The ultimate goal of this study is to develop the mMET into an English learning tool with two functions: (i) a test that can assess English proficiency with less effort, and (ii) a resource that allows enjoyable English learning while maintaining high learner motivation. As a preliminary investigation, this paper reports on an experimental implementation of the mMET using dialogue audio from
The Wizard of Oz (
Fleming et al., 1939), as well as the findings derived from this study.
The structure of this paper is as follows. Section 2 introduces the basic characteristics of the MET, highlighting both its strengths and remaining challenges. Section 3 presents the movie used in the mMET, The Wizard of Oz, and discusses its level of English and notable audio features. This section also provides detailed information about the specific movie dialogues extracted for the mMET, including the total length of the audio, the number and roles of speakers, and the auditory factors that may contribute to listening difficulty. Section 4 reports the results of a statistical analysis examining whether there is a significant correlation between scores on the mMET and those on other reliable English proficiency tests—namely, the Common Test (the current university entrance examination in Japan) and TOEIC. Section 5 discusses the implications of these correlation coefficients, exploring how movies might be more effectively utilized in English language education. Finally, Section 6 provides a summary of the study.
II. PREVIOUS RESEARCH
1. Basic Characteristics of the MET
The MET developed by
Maki et al. (2003) has the following characteristics. Participants listen to approximately five minutes of English audio and fill in about 60 to 70 blanks on an A4 sheet with appropriate English words (no more than four letters each). They reported a relatively strong correlation (
r = .68) between the MET scores and the scores of the English section of the Center Test conducted in January 2002. In their MET, the English text and corresponding audio used are taken from
This is Media.com (
Kawana & Walker, 2002), an English textbook for first-year university students. Details are shown in
Figure 1.
As shown in
Figure 1, the original version of the MET by
Maki et al. (2003) is a test in which test takers fill in 72 blanks within an English passage printed on a single A4 sheet, while listening to an audio recording. Each blank is to be filled with an English word of up to four letters, such as
life and
dog. The audio is read aloud at a pace of approximately 125 words per minute and lasts for about five minutes.
Subsequently, the Maki group continued to revise the test and develop the MET using various English passages (
Maki, 2018). The correlation coefficients between MET scores and scores on the Center Test, administered from 2002 to 2009, ranged from
r = .53 to .72 (e.g.,
Goto et al., 2010). Furthermore,
Maki et al. (2012) created a revised version of the MET using the same passage as the original version, but with blanks appearing every six words.
2 In many later versions of the MET, a rule of placing blanks at nearly regular intervals has been adopted and the English audio materials have undergone continual refinement, including the introduction of multi-speaker dialogues and versions featuring English songs (e.g.,
Wang, 2022).
2. Research Target of the MET
One of the major strengths of the MET as an English proficiency assessment is its ability to provide an approximate prediction of scores on other reliable English tests in just five minutes. However, there remains room for improvement if the MET is to become a more beneficial tool for ESL learners. The recordings used in the MET are typically spoken by professional native English speakers, presented without any background noise, and delivered fluently and clearly at a consistent speed. This style of audio is not uncommon in English learning materials. Previous research has pointed out that listening materials for English learners tend to be designed to be “listener friendly.” For example, according to
Porter and Roberts (1981), such materials are often characterized by: (i) the absence of background noise, (ii) clear intonation and pronunciation, (iii) relatively slow speech rate, and (iv) a sequence of complete sentences without omission or reduction of expressions.
Such listening materials, however, differ significantly from the English spoken in real-life contexts. In reality, it is quite rare for conversations to occur in environments entirely free of background noise. Most of our daily interactions take place amid various types of ambient sounds, such as the noise of vehicles passing by, conversations of people nearby, or background music in restaurants. Moreover, the pace of natural conversation is rarely consistent. Speakers may pause for extended periods when they forget what to say or when they hesitate; they may also speak rapidly in moments of urgency or even yawn mid-sentence. Additionally, the environment in which conversations occur can greatly affect intelligibility. For instance, speech may echo in tunnels, and voices may become muffled when someone is speaking while eating.
For ESL learners, such speech may often be perceived as “difficult to understand.” Nonetheless, it is precisely these types of acoustically “burdened” or imperfect conversations that more accurately reflect real-world English communication. Therefore, if ESL learners are consistently exposed only to artificially enhanced, easily intelligible audio materials, without being trained to process more realistic speech, it is unlikely that their practical communicative competence in English will be adequately developed.
3. Proposal of the MET
Building upon the preceding discussion, this study argues that in order to enhance ESL learners’ practical English proficiency, it is desirable to develop a version of the MET that incorporates conversational audio containing various elements of “reduced intelligibility” and more closely resembles real-world spoken English. To this end, the present study proposes the use of conversational excerpts from movies as suitable audio material for constructing such a revised version of the MET.
There are several reasons why the use of conversational audio from movies is considered desirable in the development of a more realistic version of the MET. First, movies provide a wide variety of background noise environments, allowing for the simulation of diverse conversational settings. These include conversations that take place over the phone, inside moving trains, or on beaches with the sound of waves in the background. Such variability mirrors the acoustic conditions of real-world communication. Second, emotional expression is frequently observed in movie dialogues. Many scenes involve characters speaking while displaying strong emotions, such as during drunken exchanges, intense arguments, or conversations while crying. These situations are representative of interactions that often occur in everyday life. Third, many movie conversations show variations in speaking speed. For example, angry speakers often talk quickly, while those who are intoxicated or sleep deprived may speak slowly. Learners must be able to understand such irregular speech rates in real English conversations. In addition to these features, movies also portray realistic conversational dynamics, such as sudden changes in the number of participants, interruptions by others, and abrupt shifts in topic. These characteristics make movie dialogue a valuable source for exposing ESL learners to the complexity and unpredictability of real-life spoken interaction.
In fact, previous studies have pointed out that the audio in movies possesses characteristics that differ from those found in conventional English learning materials.
Fujita (2017a) argues that while readability and vocabulary levels are comparable between textbooks and movies, movies exhibit faster speech rates and shorter or fewer pauses between sentences. Furthermore,
Fujita (2017b) reports that ESL learners performed significantly worse on a dictation test when using movie audio, with participants commenting that they had more difficulty understanding the content of the movie. These findings suggest that exclusive exposure to scripted audio designed for educational purposes may leave the ESL learners unprepared for real-life communication in English. Accordingly, this study argues for the necessity of incorporating movie audio into English language learning to foster more practical and authentic communicative competence.
III. METHOD
In order to propose a novel application of movies in English language education, this study introduces the development of an mMET, which uses audio extracted from movies. The materials aim to help ESL learners acquire practical language skills while providing an enjoyable learning experience that maintains high learner motivation. The specific rules for this approach are outlined below.
1) mMETs are created from public domain movies.
2) Each version of mMET is limited to a maximum of five minutes.
3) Target words for the blanks must be four letters or fewer.
1. Material
In this study, two versions of mMET were created using dialogue from
The Wizard of Oz. Among films in the public domain, this particular movie was selected as the source material for this preliminary study for several reasons. First, although it is an old movie, it is a highly recognizable work, which allows ESL learners to feel a sense of familiarity and continue their English learning even after the test. Second, the Japanese publishing company FOURIN has released a series of books called SCREENPLAY, which provide transcripts and English explanations for each scene of the movie. Therefore, Japanese ESL learners who refer to the SCREENPLAY edition of
The Wizard of Oz can review the test-related English expressions in detail and further develop their learning by studying English used in other scenes. In both the book (
Soneda, 2012) and on the company’s website, the educational value of this movie as a resource for English language learning is described as follows:
3
1) The English used in this movie reflects the language that native American speakers have been exposed to since childhood.
2) It contains a wide range of expressions that are applicable in everyday conversation, and many of its grammatical constructions and idiomatic phrases appear in standard dictionary forms.
3) As such, it serves as an accessible and appropriate learning resource for individuals who may feel less confident in English but wish to begin studying the language.
In addition, the English spoken in the movie is characterized by the following specific features:
1) Most of the characters speak standard English with little to no regional accent.
2) Dorothy, in particular, speaks slowly and clearly, making her lines easy to understand.
3) Because it is a musical, the film provides an enjoyable learning experience, especially for beginners.
Based on the aforementioned characteristics, in the SCREENPLAY series, the English used in
The Wizard of Oz is classified as beginner level among the categories of beginner, intermediate, and advanced. As a preliminary study, this research extracted two scenes from
The Wizard of Oz and developed mMET 01 and mMET 02 based on the dialogues. The content of mMET 01 is presented in
Figure 2 below:
mMET 01 consists of 44 blanks, and the corresponding audio lasts for 3 minutes and 2 seconds. The conversation involves three female speakers. The reason this scene was selected for the mMET 01 is that the conversation exhibits the following characteristics: (i) there are occasional pauses during the dialogue, (ii) loud sound effects occur from time to time, (iii) the Witch speaks in a hoarse and rapid voice, and (iv) the Witch disappears in the latter part of the conversation. These characteristics are rarely found together in the audio materials of conventional English textbooks, which makes listening tasks more challenging. However, although uncommon in English learning materials, they reflect a type of conversation that frequently occurs in real-life situations.
Next, the details of mMET 02 are presented in
Figure 3.
mMET 02 consists of 30 blanks, and the corresponding audio lasts for 3 minutes and 46 seconds. The conversation involves four male speakers and one female speaker. The main features of the dialogue are as follows: (i) intermittent loud sound effects, and (ii) instances where the male speakers speak in frightened voices, while the female speaker speaks in an angry tone. Conversations that are strongly influenced by psychological or emotional factors are rare in the audio materials of conventional English textbooks and can be considered a distinctive feature of movie audio.
2. Participants
Using mMET 01 and mMET 02, surveys were administered on three occasions—June 2021, October 2021, and April 2022. To examine the correlation between mMET scores and those from other established English proficiency tests, participants were recruited from university students who had taken the Common Test (the CT, hereafter) in January of the same year they sat for the mMET. The first two surveys were conducted at a university in Japan’s Tokai region, and the final one at the same university as well as another university in the Kansai region. Consequently, most participants were first-year university students. Before taking the mMET, participants received the following instructions, consistent with those used for the standard MET:
Instructions
1) Write your score on the English section of the CT taken in the same year, as well as any other English proficiency test scores (e.g., TOEIC), if available.
2) Listen to the movie dialogue and fill in each blank with a single English word.
3) The audio lasts approximately five minutes.
IV. RESULTS
In this study, the scores from two types of the mMETs were analyzed using simple linear regression to examine their relationship with scores on the CT and TOEIC. In line with previous analyses of other METs, this study follows
Yanai (1998) for the interpretation of correlation coefficients. The correspondence between correlation coefficient values and their descriptive characteristics is shown in
Table 1.
Then, the results of the mMET conducted in June 2021 are shown in
Table 2. Although the number of participants was relatively small (22 for the CT and 19 for TOEIC), statistically significant correlations were found between the mMET scores and both the CT and TOEIC scores. In particular, mMET 01 showed a strong correlation with TOEIC scores, with a correlation coefficient of
r < .83.
Next,
Table 3 presents the results of the mMET 01 and mMET 02 conducted in October 2021 with an increased number of participants. This survey involved approximately twice as many participants as the survey depicted in
Table 2. In both
Table 2 and
Table 3 surveys, mMET 01 and mMET 02 scores exhibited statistically significant correlations with scores on the CT and TOEIC.
Finally,
Table 4 below illustrates the findings from the mMET conducted in April 2022.
4 The results of the survey shown in
Table 4 indicate that, even with more than 150 participants, both mMET 01 and mMET 02 scores demonstrated statistically significant correlations with both the CT and TOEIC scores.
However, it should be pointed out here that in the survey shown in
Table 4, although the number of participants increased, the correlation coefficients decreased compared to those in the surveys presented in
Tables 2 and
3. This change may be attributed to several factors, such as the use of a larger room and more participants in the experiment reported in
Table 4, which may have reduced the clarity of the audio, as well as the fact that the surveys in
Table 4 were conducted by different experimenters at the two universities. If these possibilities hold true, it is evident that future development of the mMET must systematically address and mitigate such undesirable confounding factors.
V. DISCUSSION
As demonstrated in the previous section, the surveys conducted in June 2021, October 2021, and April 2022 revealed that scores on both mMET 01 and mMET 02, developed from dialogues in The Wizard of Oz, exhibited statistically significant correlations with scores on the CT and TOEIC, although there were slight differences in the strength of the correlations. Importantly, this study, which utilized audio from the movie, showed that despite incorporating factors that increase listening difficulty—namely, (i) conversations involving three or more participants, (ii) background noise such as music and sound effects, (iii) variable conversational pace, and (iv) emotionally charged voices expressing anger or fear—the mMET functions comparably to other METs as an effective measure of English proficiency. In other words, the mMET demonstrates substantial potential to serve as a novel English learning tool that employs more authentic, real-world audio.
Of course, there are already existing English learning materials that utilize audio resembling real-world conditions. For instance, Unit 7 of
Breakthrough Plus, Level 1 (
Craven, 2016), published by Mcmillan Education and designed for beginner level English learners, includes conversation audio containing various background noises. The phone conversation, conducted amid radio voices and background noise, proceeds as shown below.
Megan: Hello, Robert. It’s Megan. Are you busy?
Robert: Oh hi, Megan. I’m, er … filling out an application form, actually. I’m applying for a part-time job.
Megan: Really, That’s great. I’m really pleased you’re finally looking for work. Hey, what’s that noise? Robert: Er, I’m listening to the radio. It helps me think. Anyway, what are you doing these days? How’s the job going?
It may be asked why it is necessary to develop teaching materials using audio from movies. There are several reasons for this. First, as pointed out in previous studies—and as noted at the beginning of this paper—materials that incorporate movies have been shown to enhance the ESL learners’ motivation to study English. Second, the English audio used in the mMET is extracted from specific scenes in movies. This allows learners who are interested to review the tested content not only through the audio but also by watching the corresponding scenes with visuals after the test. Furthermore, they can continue learning by watching scenes beyond those included in the test. In contrast, audio recordings created for general teaching materials are not intended to be used beyond a few minutes—they are not designed to encourage additional learning. For instance, the phone conversation referenced in Craven (2016, p. 44) lasts only one minute and fifteen seconds. Third, movies can offer a significantly wider range of conversational settings. These may include, for example, dialogue during a violent storm, arguments between romantic partners shouting at each other, or whispered messages between spies. It is almost impossible for standard educational audio materials to reproduce such variety.
From the perspective of further developing the MET, there are also clear advantages to using English audio from movies. One notable characteristic of many existing versions of the MET—including the original version introduced in
Figure 1—is that the English passages are read at a constant speed. In such cases, it is possible to predefine the number of words and insert blanks at fixed intervals. For example,
Figure 4 below illustrates a part of a version of the MET in which every sixth word is blanked out, using the same passage as the original MET. This version is called the MET 6B in
Maki et al. (2012).
In audio used for traditional METs where the script is read at a consistent speed, such as in the MET 6B, it is possible to insert blanks at regular intervals. However, it is difficult to place multiple blanks in adjacent word positions. This is because test takers would not have enough time to write down the missing words while listening to the audio. In contrast, the mMET includes many audio passages in which the speaking rate is not uniform. In these cases, such as conversations that contain long pauses or lines delivered intentionally at a very slow pace, it becomes feasible to place blanks for two (or more) consecutive words. The pauses and the slow speaking allow test takers sufficient time to write down more than one word. This means that it may be possible to design a version of the MET that maintains the conventional number of blanks (approximately 60 to 70) but can be completed in less time than the typical five minutes required by the traditional MET format.
Therefore, the development of the mMET using movie audio has the potential not only to produce learning materials that are more enjoyable for ESL learners and better suited to fostering practical language skills, but also to create tests capable of assessing learners’ English proficiency in a shorter amount of time. In this paper, the analysis was limited to two versions of the mMET constructed using audio from The Wizard of Oz. However, it is evident that future work should involve developing mMETs based on a broader range of movies as well as selected individual scenes.
VI. CONCLUSION
This paper reported the results of a preliminary study on the development of a version of the MET that uses movie dialogue, specifically by creating two versions of the mMET based on conversations from The Wizard of Oz. The scores on these mMETs showed statistically significant correlations with those from other reliable English proficiency tests, such as the CT and TOEIC. However, the correlation coefficients varied considerably, ranging from r = .46 to .83. This variation suggests the need for further investigation into the contributing factors, and it goes without saying that continued efforts are required to develop more precise versions of the mMET.
Although the development of the mMET is still in its early stages, the project holds considerable promise. From the perspective of MET development, this study represents the first attempt to create the MET using audio from movies. From the perspective of the use of movies in English language education, it is also the first attempt to use movie audio to predict scores on other standardized English proficiency tests. For these reasons, there remain numerous issues to be addressed in the continued development of the mMET.
For instance, in The Wizard of Oz version presented in this paper, the transcript layout was designed so that each new speaker’s utterance begins at the left margin. As a result, a large amount of unused white space appeared on the right side of the A4 page. Moreover, this layout may place an unnecessary burden on test takers, as they are required to repeatedly shift their gaze back to the left margin while listening to the audio. Therefore, future work must carefully consider more effective layout designs.
In addition, adjusting the selection of blanked words by part of speech could help identify which grammatical categories contribute most to the ESL learners’ comprehension. Varying the types of background noise inserted into the dialogue may also provide insights into which auditory elements most affect the ESL learners’ listening performance. These topics warrant further investigation, which will be explored in future research.