Showing posts with label listening/speaking. Show all posts
Showing posts with label listening/speaking. Show all posts

Wednesday, November 7, 2007

Teacher Verification of iBT

While in line for tickets for the university's production of Little Women the Musical (we're going on my birthday for a pre-Thanksgiving holiday warm-up), I finished another integrated tasks evaluation study:

Cumming, A., Grant, L., Mulcahy-Ernt, P., & Powers, D. (2004). A teacher verfication study of speaking and writing prototype tasks for a new TOEFL. Language Testing, 21, 107-145.

You may notice that this is an article written by Alister Cumming who also wrote the iBT integrated tasks textual analysis article that I also read. I also referred to Cumming's other writing test related work for my MA thesis as well as my rater decision-making study (which I need to finish up and then develop into an article). Cumming seems like a very busy man, but he has been kind enough to respond to a couple of emails that I have sent him about his research and about the UToronto program.

Summary: This article attempts to provide content, context, and concurrent-related validity evidence in regards to the speaking and writing tasks for the new TOEFL (which is now known as iBT). As with the previous studies that I have reviewed, this one focuses on both integrated tasks (combining writing/speaking with listening or reading) and independent ones (in which examinees use their personal experience or opinions to complete the speaking/writing tasks).

Whereas the previous articles focused on the language content (Cumming, 2005) and the scoring procedures (Lee, 2006), this one focuses on how teachers of ESL students feel about the structure of the iBT tasks and their students' performance on these prototype items.

Primary questions are:
  • Is the content domain of integrated tasks perceived to correspond to the demands of academic English requirements of university study?
  • Is the performance of examinees on these tasks perceived to be consistent with their classroom performance?
  • Are the tasks perceived to be adequate evidence for making decisions about examinee language ability?
The researchers found that the teachers felt that these tasks were fairly authentic and represented a variety of language skills required at university (so far as a test is able to recreate authentic situations). There were some concerns that the tasks were not fair for lower ability students who performed poorly on integrated tasks if they did not understand the input material; however, for the most part, teachers felt that student performance on the tasks were indicative of classroom ability. The greatest concerns that raters had with the evidence claims of the tasks were that some students may do poorly if they feel uncomfortable sharing personal opinion (for the independent tasks) or if they struggle with the stimulus content (for the integrated tasks). Cumming et al. conclude with some suggestions for improving the task format and content. They also suggest that students would benefit from exam preparation in order to understand the rhetorical and educational aims of the exam. Lastly, they suggest that more research be done into standardized ratings and on teacher perceptions of students' ability in the classroom versus an exam.

Critique: It's a shame that this article it stuck in the middle between an extensive evaluation (involving a variety of teachers and locations) and an intimate study (focusing on one or two individuals and their unique situation). As a result, we have neither the generalizablity of a large statistical study nor the valuable insight into a specific context. Instead we end up with a bunch of incomplete thoughts and ideas. It's as if Cumming et al. scratch the surface of several interesting ideas, but never uncover a single one. Still, for all its limitations, this study provided me with a good model for some qualitative research (including a survey that I adapted) and it is also likely to help me form some of my content and context validity questions. I also decided, as I read its account of teacher-based evaluation, that a student-based evaluation would be a great compliment to such a study. We'll see if I can make it work, or if I will have bitten off more than I can chew.

Connections: This study served to show me that I need to read a lot more about teacher evaluation study about tests. Was this a good example? How have others approached this avenue of validation? I did enjoy seeing how this person-focused (as opposed to text-focused or score-focused) study complimented other, more quantitative, studies on the iBT in order to give a more complete and well-rounded assessment of the exam.

Additional Reflections:
  • I would really like to do an evaluation of the integrated writing tasks at the ELC that incorporates myriad sources of validty evidence including teacher-evaluation, student-evaluation, textual analysis, score-analysis, and possibly more. Am I crazy to want to do this much? Will TREC approve such a study? Will I find a dissertation committee who will approve such a multi-faceted approach?
  • Who will I choose as teachers for my teacher evaluation? I think that I would prefer to use all of them (provided that they agree to participate in the study). I don't want to exclude someone because their reasons for not being eager to volunteer may actually be connected to their feelings about the exam and are therefore a real concern that I need to take into account.
  • Will I have time to do this for speaking as well as writing? My tentative schedule, which I need to present to potential committee members next week, outlines an evaluation of writing for winter 2008 and speaking in summer 2008. I think I can be ready, but the issue is as much about analyzing data as it is about collecting it. Who will I get to help, especially when I am so busy around exam time with administrative duties?
  • What kind of concurrent validity claims might I make? Will I ask teachers to rate their students and then compare performance with expectations, or will I use data that we already collect (classroom scores or rated evaluations)? Are teachers very good at predicting student ability? Lee's 2005 thesis of the L/S tests at the ELC say no (as did major portions of this study), but I think that both of those are problematic: Lee because the teacher ratings were probably based on Speaking when she was comparing those to Listening scores, and in Cumming's case, he admits that BICS/CALP may have a lot to do with false impressions. Of course BICS/CALP may be an issue for me too, but if students and teachers practice items in the classroom more than once, then teachers ought to be good at predicting success, shouldn't they?
  • How will tasks be adjusted for lower levels (adhering the concerns of Sara in this study)? Laura (our Level 1 teacher) has already told me that she thinks Level 1 students need to be able to hear the integrated listening passage twice. As it is, they barely understand it before it's over. A second listening for Levels 1 and 2 might be appropriate and justified (given the findings of this study and Laura's experience).

Tuesday, November 6, 2007

G-Study of iBT

The most recent paper that I finished reading is:

Lee, Y. (2006). Dependability of scores for a new ESL speaking assessment consisting of integrated and independent tasks. Language Testing, 23, 131-166.

Summary: Lee conducts a generalizability study of prototype speaking tasks for the iBT (internet-based TOEFL). The purpose of the study was to assess to what degree reliability increased with the number of tasks and ratings.

Tasks include three types: integrated reading-speaking, integrated listening-speaking, and independent. The effect of rater was ignored due to the difficulties of orchestrating a fully-crossed design; however, rating was considered as a facet, so all ratings were analyzed in a double-rating model.

Lee concludes that increasing tasks has a greater influence on increasing reliability than does increasing raters. Lee also found that there was sufficient evidence to justify collapsing analytic scores into a holistic score of speaking given that correlations among task types were quite high.

This has implications for the iBT speaking component. First, including more than one speaking task is more likely to increase the reliability of scores than adding extra raters. In other words, Lee suggests that more reliable scores are given when raters have a greater number of speaking sample from the examinees than when multiple raters are used. In essence, many raters can still come up with divergent scores when only one sample is used, but when many samples are provided, even a small number of raters is highly accurate at scoring the examinee.

The second implication affects the use of holistic scoring. Rather than require raters to score each sample separately (and then averaging the separate scores into a composite score for each examinee), Lee suggests that since all task types tend to correlate highly, there is sufficient evidence to warrant the use of a holistic score for each examine based on an overall impression rather than a combined average. The benefit is that this reduces the time it takes raters to score a single examinee.

Critique: Admittedly I am not impressed that Lee did not attempt a fully-crossed design. I think this makes the "rater has low effect in increasing reliability" claim given that we have nothing more than a 1-rating versus 2-ratings option. I recognize that fully-crossed designs are really hard to manage, but Lee was funded by ETS (who creates the TOEFL), so I don't know why ETS didn't demand a more rigorous research design.

I am also concerned that Lee does not describe the testing occasions. Did students take all 12 tasks at once, leading to enormous test fatigue? Or were the testing occasions spread out over several weeks thereby possibly confounding test results given that students may have improved in proficiency over the period of testing. I need to email Lee and ask about this specifically.

Connection: Thankfully I have been introduced to generalizability theory in LING 660 (Language Testing) and IP&T 752 (Measurement Theory), because it is hard for a researcher to really explain it well in a short article. I had also previously read a study (for my thesis) written by Rob Schoonen (who is a much better writer than Lee) and this helped me to understand the analysis and results that Lee described.

Lee does a decent job of connecting the results of this study to others. I found myself coming to similar conclusions and making connections to my limited experience with G-studies (such as Schoonen). I also found it valuable to read this study in connection to the iBT validation studies by Cumming. It helped me gain a bigger picture of the iBT validation project and how all these different elements function together to inform the development of this high-stakes exam.

Additional Reflections: This article sparked a lot of ideas for me.
  • Could my study include a comparison of analytic/holistic scoring? Or at the least, could I compare portfolio scores (holistic) with scores for just the integrated tasks? This may justify the need for a separate score for draft writing and timed writing.
  • I like that Lee suggests that test developers need to state the purpose of the task. Are independent tasks just about language, and integrated tasks just about content? In our case, probably not. They are both about language. This need to be clarified.
  • A scoring-related validity component would be a valuable part of my evaluation that might also include teacher and student content/context validity evidence.

Tuesday, October 9, 2007

Moving forward

I heard back from IRB today. My request to extend my research has been granted. Of course now that I want to do a different (but related) study for my dissertation, I will need to submit a new IRB application.

Why start something new for my dissertation? I'm applying for assistant professor positions (to start next Fall) and one of them is looking for expertise in an area that I haven't focused on very much: oral assessment and listening/speaking pedagogy. Of course I had taught L/S classes, and I'm in the process of becoming OPI certified, but I haven't done much on this topic from a research perspective. So now's my chance.

Truth is, whatever an employer wants me to be an expert in, I can do it. I have only become focused on writing assessment and methodology because:

1) My thesis chair wanted to research it, so I did it for my thesis, and
2) My current job as the writing coordinator requires me to become an expert on it.

So if a potential position requires expertise in something else, I can do it. My revised dissertation project will:

1) Build on my current research,
2) Relate directly to what I do for my job anyway, and
3) Extend my expertise to include research on L/S assessment.

So we'll see how it goes. The fact is, if I want to be ready for these jobs that start Fall 2008, then I need to be all but done my research by then, which gives me less than a year to plan, conduct, and write my dissertation. Heh. Who knows? I could do it. Probably. Maybe. Yeah. And if nothing else, this job application process will motivate me to graduate ahead of schedule. And there's nothing bad about that idea.

Tuesday, August 7, 2007

Video Creation as a Learning Tool

When I was an ESL teaching assistant in Hawai'i, one of the major assignments that I helped with was a Listening/Speaking class video project. This one class of beginning level students were divided into groups and had to plan, write, act, film, and edit short films. They had a blast doing it and their movie sharing night was a big hit with them and all of their friends.

It had always been my intention to do a similar project in the future, but it just never seemed to work into the curriculum. Until this week. This summer I have been teaching a beginning level Listening/Speaking class and we finished with our required tasks last week, so either we could spend a whole week reviewing grammar or listening to sample dialogues (in preparation for their final on Thursday), or we could do something productive.

I proposed the film idea to the class yesterday and they jumped at the opportunity. I was a little concerned about the feasibility of planning, practicing, and filming a movie in 2 days, but I figured we would at least give it a shot. Worst case scenario: we don't make a film, but they get lots of practice reviewing the phrases and tasks for their final exam.

So when I arrived to class this afternoon and found the chalkboard filled with their storyboarding, I was impressed. They managed to find a story that used a wide variety of their L/S tasks while involving all the class members into a naturally progressing and cohesive story. We already film the major scene today, and we will finish it up tomorrow. I'll edit the whole thing over the break, and their movie will become one of the samples for the *ALL NEW* ELC Film Festival Night 2007 (a new activity that we are hoping to try out this fall).

If the project works out, I will suggest the idea to other L/S classes. It's a motivating project and it encourages students to use productive language in near-authentic situations.