Wednesday, October 15, 2008

New blog

Just in case anyone follows a link to this blog...

I now post my academic research thoughts at http://robblogva.wordpress.com/

It's pretty dry, unless you are a language testing researcher, in which case it is still fairly dry.

Sunday, March 30, 2008

Next Step

As it turns out, I'm gaining more job interviewing experience than I anticipated. One of my applications has turned into a telephone interview, followed by a request for an in-person interview next month. The job is with a university in another state, so this means that I'll get to visit a new part of the country.

I wish that I had more experience as a job interviewee, but in truth I've only ever interviewed for a handful of jobs - and three of those were for positions with my current university. I think that during interviews I speak too quickly, I don't answer questions directly enough, and I forget to ask the important questions that I want to know. But I think that I can get better. I've started prepare for interviews more effectively by anticipating general question topics and articulating in writing (before the interview) my response. I've also started to write out my own questions so that I can make sure that I ask them during the interview instead of kicking myself afterwards for forgetting.

And whether I get a job offer or not, I will feel far more prepared for the next time I interview. Thankfully I still have 16 months left on my current contract, and I'm excited about my new job duties, so whether I stay or whether I go, I feel very fortunate to have good job prospects.

Monday, March 24, 2008

Getting Job Ready

We recently went through our interviewing and hiring process at work. We are looking to fill two positions that will become vacant when two of my co-workers complete their 3-year non-renewable contracts at the end of this summer.

We tried a new interviewing process which went really well: rather than interview each candidate for 45-60 minutes in a room full of 15 or so interviewers, we decided to break the interview into 3 smaller, more focused interview with smaller interviewer groups. In each room, the candidate was asked questions and invited to participate in a role play that focused on a different aspect of our institutes's mission statement. Although not a perfect system, it's a big improvement over the old process, and I think that the candidates would agree (in fact one of them did who had experienced the previous process the year before).

I am also subjecting myself to the interview/job application experience. Although my contract is still good for another 18 months or so, I decided that I would start gaining experience in the job market so that I find something before I'm out of a job.

I began applying to jobs outside of my experience range, just to see what the process worked. It was great for me to update my CV and prepare application letters and other associated documentation. About a month ago, responses starting to come in: Thank you, but no thanks. These weren't shocking responses, but I admit that it's still a little disappointing to have gotten 3 rejections without so much as an interview.

Of course I did apply to a few positions that are more in my range. Although most any job that I applied for would really like more experience (not to mention a completed PhD) than I have, I still have several more than have not sent a response. And just last week I telephone interviewed for another one with a university in a wonderful, green, college town.

I don't know that anything will come of these applications, but it's great to keep my CV current so that when the time does come that I need to begin a serious job hunt, I will know that I have most of my work already done. And hopefully after a few more interviews, I won't be so nervous when the right job does come my way.

Friday, January 18, 2008

JITT in IP&T

This is week two of reading the JITT (Just in Time Teaching) quiz results from my students.

Wow! I really like this approach to class preparation and lesson planning. After reading the JITT quizzes responses, I feel as though I have already held a mini-class with my students and I have an idea of which points from the reading are confusing for them, which are interesting to them, and which are clearly understood by them.

How does it work?

1. Kimberly (my wife and fellow instructor in this type of course) writes some quiz questions that are designed to (a) highlight the vocabulary that the program director feels are most important and (b) get students engaging with the most valuable ideas from each chapter. Three of the items are fixed-response (multiple choice or fill-in-the-blank, etc.) and one is usually an open-ended application question. We have a mis for obvious reasons: multiple choice are fast for students to do and cover a wide range of material quickly, but the open-ended questions lets us see how students problem solve and whether they understand the material at a deeper level.

2. Kimberly send me the questions and I review them, providing editing, content, and test-quality feedback. We make some revisions through discussion.

3. We post our questions online using the university's course management system (CMS) which has a quiz feature (a little cumbersome, but it works).

4. Our students are required to read the material and then answer the questions by a deadline (I have them complete the quizzes by midnight of the day before the day we meet in class), so that as instructors we then have time to analyze their responses and develop our lesson plan.

5. We download the quiz results. At first, we did this manually: we openned up every single response from every single student, and recorded the information from a web browser into an speadsheet file. This was horribly, painsakingly long (especially given the cumbersome nature of the CMS package). Of course after doing this the first week, I noticed that the CMS has a "Download Quiz Results" link. Needless to say, we tried this for week two and it was a huge timesaver.

6. We analyze the results. This is not a complicated process. Really, we are interested in seeing which questions (and hence which concepts) the students have a clear understanding of, and which concepts are problematic. Generally, I don't write any notes for questions in which all students got correct (or only one student got wrong - instead I would suggest talking to or emailing that student directly). However, in questions where there appears to be various different responses, I take notes on the most common misconceptions.

In the case of open-ended questions, students tend to respond with the most common, salient features from the text (most often the ideas that are highlighted in the end-of-chapter summary page) but few demonstrate understanding of the fullness of the concept as expressed in the text. I make notes of these deficiencies and I highlight responses from students whose answers indicate that they have grasped these finer points.

There is one final question that we analyze: the optional, open-ended feedback question. At the end of each quiz we include a feedback field for students to tell us how the reading is, any challenges that faced, and any other thoughts that they have about the reading, the quiz, or the class. At least half of the students skip this optional question, but the feedback from those who do respond is invaluable. Frequently students will explain how they had trouble with a multiple choice item ("I felt like there were two correct choices because...") or they will express concerns about a particular topic ("I'm still kind of fuzzy on the idea of..."). And about half of the responses are just postive feedback ("I really enjoyed this chapter since I learned something new that I can apply to..." or "I found this chapter much easier to understand than the last since I had already studied a book about..." or "Question 3 was a really good measure of the chapter since it pulled together all of the ideas into one practical situation.").

These notes (list of what students do not understand as well as their additional concerns and questions) serve as the guide for my class presentation.

7. We write a lesson plan. Although our class meets for a 3-hour block each week, we generally only try to take about 1 hour for our JITT presentation since we reserve most of our class period to student presentations (especially important since we teach pre-service teachers who need the practice of planning and presenting mini-lessons). So we design a presentation that covers the most difficult chapter concepts based on the JITT quizzes.

It's only week two, so I am hardly an expert at this. And I rely heavily on Kimberly when it comes to desiging a JITT presentation (aka lecture) since she has done an extensive review of the literature for one of her research projects. I really appreciate being able to work together on this experiment and I hope that Kimberly and I will be able to contribute to the JITT field by writing a paper that describes how to move from JITT quizzes to JITT lectures - a part of the JITT process that seems to lack publications.

Thursday, January 10, 2008

My English Teacher, the Anti-Christ

Before a new semester gets underway, I thought that I would share this story from last semester.

Due to a teacher shortage, I got stuck with a double-section of our upper-level writing class. Normal class sizes were maxed out to the highest that I had seen in our Center in the few years that I worked there, and here I ended up with a double-section. Although I had taught larger class sizes than this before (in China), I had never taught writing to such a large group. I thought I was going to fail miserably.

I suppose that it did no help that I am trying to build a reputation as the meanest teacher in the whole Center (fueled only by my pronouncements to my class that I am the meanest teacher in the whole Center). It is possible that some students actually believe this declaration, but as the semester wore on, it became apparent that few (if any) put an credence in this statement. However, by wonders of wonders, I may have actually convinced a single student.

In my course evaluations, I got the usual teacher feedback (helpful, prepared, organized) along with the usual course feedback (relevant material, writing is important but hard, good preparation for university, I think this class is too big). So no surprises, but I did enjoy the comments from one student:

This teacher has no Christ-like attributes. He is no caring, kind, patient. He is cold-hearted person.

Not to make light of this student's apparent pain (nor to make fun of his grammar), but I admit that I smiled - and even chuckled - when I read this. As I completed reading the evaluations, another interesting comment came up.

On the first major paper the teacher gave me zero score and I was totally shocked and upset. Can you believe it? This is horrible teaching style.

Anyone want to wager a guess that the first and second comments came from the same student? Course evaluations are supposed to be anonymous, but when a student makes revealing comments such as these, it doesn't take much detective work to figure out that these comments were made by the student whom I warned about plagiarism in his draft (and who sent me a rude email explaining that he had never been so insulted in his life and to whom I responded calming and kindly even though I was more than a little angry at his rude email and whom I invited to meet with me but who never came and who then quickly shut up when I explicitly and privately showed him the plagiarism in class), and then who did little to fix the plagiarism and as a result got a zero score just like I had warned him repeatedly.

So what do we learn from this story?

If you persist with plagiarism, your teacher is the Anti-Christ.

I think that is a good moral. Can't wait for a new semester of writing students. I'll share this story and maybe, finally, students will believe it when I tell them that I am the meanest teacher in the Center. After all, I have a student's testimony.

Monday, January 7, 2008

A break from taking a break

I used to be a regular blogger, but perhaps the challenge of full-time work and full-time studies has overridden my blogging discipline. Or maybe I've nothing new to blog about since I essentially work in the same place I did when I was an MA student, and I have gone back to school taking similar kinds of classes that I did when I was an MA student. Academically, I might not have anything new to add right now.

However, this new semester, I am teaching a new class: educational psychology. I am certainly not an expert in this field, but I do have a lot of experience with teaching, and I have taken classes in educational psychology, and I did get hired for the job, after all. So I guess that I'll be fine. It helps that Kimberly and I will both be teaching sections of this course, so we share ideas and plan together. I'm sure that if I get stuck with something, she'll help me out.

I'm also continuing to send off job applications. Although I do not intend to get hired this year, the process of applying has helped me to update my CV, find references, and, most recently, to write a statement of teaching philosophy. I had written a teaching philosophy a few years ago when I first started grad school (as an assignment for a course), but I misplaced it, and I figured it was time to reflect on these past few years of teaching and how I feel about it. It's hard to put all of that experience down on a couple pages of writing, but I suppose I captured the main essence of how I feel about teaching. In any case, it has helped me to self-reflect as I start a new semester with new groups of students.

Tuesday, November 27, 2007

Responding to the counteragument using academic templates

"I know what to do, but I just don't know how to do it." A common frustration in any writing class.

In these last few weeks of the semester, my upper-level writing class and I writing argumentative research papers. This assignment is pushing them to exercise new skills such as developing their own organizational outline, and encouraging them to respond to critics of their own positions.

I say "we" because I am doing this along with my students. In conjunction with the L2WRG (Second Language Writing Research Group), a few other researchers and I are "responding" to a recent debate in error correction. As my class and I discuss how to respond to research, it means as much to them as it does to me.

So how does an instructor teach how to write an argumentative research paper? Good question, and I wish I knew; this university - for all the emphasis that it places on writing - offers few courses on the teaching of writing. But despite my lack of training, I have done my best to pick up writing-pedagogy professional development opportunities. One such opportunity was a Writing Matters Seminar (sponsored by the university writing initiative) with a guest speaker who recently wrote the "They Say, I Say" composition help book.

"They Say, I Say" is based on the premise that new university students need to learn the language of academic English. Rather than simply expect students to learn this language implicitly, the authors suggest that university instructors need to raise students' awareness of these phrases and forms in order to facilitate this type of "language acquisition."

Although this book is intended for native speakers of English, it is even more relevant in our ESL class. If natiev speakers struggle to know how to put academic ideas into academic words,
then it is even more imperative that ESL students be explicitly taught these phrases of academic language.

The book offers templates for phrases and their functions, but the book does not explain how to teach them. So I'm trying to figure this out.

There are a few techniques I have used to teach phrases/vocabulary.
  1. Allow students play vocabulary games with partners using the AWL (Academic Word List)
  2. Require the use of AWL words in their essays (a portion of their grade is tied to this)
  3. Model, show examples, and encourage practice with academic phrases in a lecture format
I wish there was more. I hope to improve my teaching techniques in the near future.

Monday, November 19, 2007

Vocabulary Measures

Because I plan to have some type of linguistic measurement as part of my validation study, I have been reading about types of language statistics. The most recent is a study by Batia Laufer and Paul Nation.

Laufer, B., & Nation, P. (1995). Vocabulary size and use: Lexical richness in L2 written production. Applied Linguistics, 16, 307-322.

I am not familiar with Laufer, but I most certainly have heard about Nation. He is the man behind the Academic Word List (AWL), and list of commonly used words in university writing. We use the AWL at our center to help our students improve their reading and writing of academic texts.

This study appears to be pre-AWL since Laufer and Nation make no reference to it, and instead refer to the general service list and other university word lists. I imagine that the AWL was created not long after this study, since the inklings of the AWL are emanating from this work.

The primary focus of the study is an argument in favor of a new kind of vocabulary measure: the Lexical Frequency Profile. Although I don't exactly buy the LFP bid (the explanation of its use was verbose and confusing), I did appreciate the succinct discussion of existing measures of vocabulary use. In a matter of about 2 pages, Laufer and Nation clearly explain the pros and cons to several commonly used lexical measures. Even if I didn't get much out of the rest of the paper, I enjoyed their assessment of these measures.

The question is raises for me is: what linguistic measures will I use in my study? Almost certainly I would like to have at least one lexical measure, and hopefully this study will help me to explain and justify my decision.

Tuesday, November 13, 2007

Research Proposal Draft 1

I finished a draft of my dissertation research proposal yesterday and took it by the department to show it to potential committee members. I was well received and everyone agreed to look it over with the promise that I would check back with them later in the week. And that started today.

I met with the potential chair of the committee, whose first comment was, “Do you remember what advice I give to anyone who attempts to do a validity study for a dissertation?”

Of course I knew. The answer is, “Don’t do it.”

“It’s not that I think that validity studies are a bad idea; they are very important and should be done,” he said. “Just not as dissertations.”

“Why exactly is that?” I asked.

“They are complicated and intensive. It’s hard to get it done.” Which I knew. I recognize it. In fact, as I explained to him, I have already done the lit review and collected the data of for a study that might work in place of this validity one. It would be an easy route. But I’ve already done it.

The integrated skills study needs to be done. It is complicated, but it’s real. I’d rather do something that is needed than something that it easy. So, yes, in other words, I am crazy and stupid.

In our conversation, this professor and I discussed my proposal and how I can limit/focus my efforts while still maintaining the multi-faceted approach of a good validation study.

The next step is to flesh out specific research questions, a solid literature review, and a clear methods section (including all proposed analyses). If I can get that done by January, I think that I will be in a good position to still do this complicated study while aiming for a realistic graduation timeline.

Saturday, November 10, 2007

Language acquisition as construction refinement

Nick Ellis, a linguist from UMich, has been on campus this week discussing his theories of second language acquisition (L2A) and I was finally able to attend a session Friday afternoon. Although I would have preferred to attend an earlier session in which he discussed the need for an Academic Phrase List (in addition to Paul Nation's Academic Word List), I found his linguistically-heavy discussion of L2A theories fairly interesting. Although I have only ever skimmed the surface of L2A theories (I took my graduate course in L2A during my first summer semester of grad school), I was still able to follow some of it while I sat near the back and worked on my laptop.

As I listened, it seemed like Ellis's theories of L2A validate my teaching approach. From what I understand, it appears that Ellis is claiming that L2A is less about learning rules than it is about noticing and refining phrases (aka constructions). This is what I have been trying to get my more advanced levels to do. They frequently ask me to give them rules about language - and I try whenever possible - but the truth is that many of their questions cannot be answered by rules. Language isn't really a collection of rules, as much as grammarians would lead us to believe. Instead, language is rules by usage, not grammar. If students want answers about complex language, a rule book will not help them.

I answer these questions with corpus research. I show students that seeing how language is used is the best way for them to build and refine their own constructions of English. The easiest way to do this is to type an example phrase into a corpus viewer (such as Mark Davies's viewer). Even better, I encourage them to pay attention to constructions as they read and listen to authentic academic material. Of course students don't like this method because it takes more work, but in truth this is how I learned academic language, and this is how all native speakers learn language: noticing and refining constructions based on reading and listening.

Yet, as I came to this conclusion, a part of me questioned whether Ellis really was supporting my theory of learning, or whether I was interpreting his lecture in order to justify my own approach. Either way, I'll share my thoughts with my students next week and see what they think.

So what was I working on during the lecture? I was trying to finish the integrated writing tasks for this semester's final exams. I had written Level 1-4 earlier in the day, but due to the second language writing research group (L2WRG) meeting I had to put Level 5 on hold. So when a group of us moved directly from L2WRG to the Ellis lecture, I opened my laptop and wrapped up. Even though there was no internet access in the basement, I had taken enough notes on the selected exam prompt topic that I was able to write the reading passage without the source text. And as I wrote, I couldn't help but realize that everything that I was writing was, in fact, a series of constructions that I had learned exactly as Ellis explained it.

Wednesday, November 7, 2007

Teacher Verification of iBT

While in line for tickets for the university's production of Little Women the Musical (we're going on my birthday for a pre-Thanksgiving holiday warm-up), I finished another integrated tasks evaluation study:

Cumming, A., Grant, L., Mulcahy-Ernt, P., & Powers, D. (2004). A teacher verfication study of speaking and writing prototype tasks for a new TOEFL. Language Testing, 21, 107-145.

You may notice that this is an article written by Alister Cumming who also wrote the iBT integrated tasks textual analysis article that I also read. I also referred to Cumming's other writing test related work for my MA thesis as well as my rater decision-making study (which I need to finish up and then develop into an article). Cumming seems like a very busy man, but he has been kind enough to respond to a couple of emails that I have sent him about his research and about the UToronto program.

Summary: This article attempts to provide content, context, and concurrent-related validity evidence in regards to the speaking and writing tasks for the new TOEFL (which is now known as iBT). As with the previous studies that I have reviewed, this one focuses on both integrated tasks (combining writing/speaking with listening or reading) and independent ones (in which examinees use their personal experience or opinions to complete the speaking/writing tasks).

Whereas the previous articles focused on the language content (Cumming, 2005) and the scoring procedures (Lee, 2006), this one focuses on how teachers of ESL students feel about the structure of the iBT tasks and their students' performance on these prototype items.

Primary questions are:
  • Is the content domain of integrated tasks perceived to correspond to the demands of academic English requirements of university study?
  • Is the performance of examinees on these tasks perceived to be consistent with their classroom performance?
  • Are the tasks perceived to be adequate evidence for making decisions about examinee language ability?
The researchers found that the teachers felt that these tasks were fairly authentic and represented a variety of language skills required at university (so far as a test is able to recreate authentic situations). There were some concerns that the tasks were not fair for lower ability students who performed poorly on integrated tasks if they did not understand the input material; however, for the most part, teachers felt that student performance on the tasks were indicative of classroom ability. The greatest concerns that raters had with the evidence claims of the tasks were that some students may do poorly if they feel uncomfortable sharing personal opinion (for the independent tasks) or if they struggle with the stimulus content (for the integrated tasks). Cumming et al. conclude with some suggestions for improving the task format and content. They also suggest that students would benefit from exam preparation in order to understand the rhetorical and educational aims of the exam. Lastly, they suggest that more research be done into standardized ratings and on teacher perceptions of students' ability in the classroom versus an exam.

Critique: It's a shame that this article it stuck in the middle between an extensive evaluation (involving a variety of teachers and locations) and an intimate study (focusing on one or two individuals and their unique situation). As a result, we have neither the generalizablity of a large statistical study nor the valuable insight into a specific context. Instead we end up with a bunch of incomplete thoughts and ideas. It's as if Cumming et al. scratch the surface of several interesting ideas, but never uncover a single one. Still, for all its limitations, this study provided me with a good model for some qualitative research (including a survey that I adapted) and it is also likely to help me form some of my content and context validity questions. I also decided, as I read its account of teacher-based evaluation, that a student-based evaluation would be a great compliment to such a study. We'll see if I can make it work, or if I will have bitten off more than I can chew.

Connections: This study served to show me that I need to read a lot more about teacher evaluation study about tests. Was this a good example? How have others approached this avenue of validation? I did enjoy seeing how this person-focused (as opposed to text-focused or score-focused) study complimented other, more quantitative, studies on the iBT in order to give a more complete and well-rounded assessment of the exam.

Additional Reflections:
  • I would really like to do an evaluation of the integrated writing tasks at the ELC that incorporates myriad sources of validty evidence including teacher-evaluation, student-evaluation, textual analysis, score-analysis, and possibly more. Am I crazy to want to do this much? Will TREC approve such a study? Will I find a dissertation committee who will approve such a multi-faceted approach?
  • Who will I choose as teachers for my teacher evaluation? I think that I would prefer to use all of them (provided that they agree to participate in the study). I don't want to exclude someone because their reasons for not being eager to volunteer may actually be connected to their feelings about the exam and are therefore a real concern that I need to take into account.
  • Will I have time to do this for speaking as well as writing? My tentative schedule, which I need to present to potential committee members next week, outlines an evaluation of writing for winter 2008 and speaking in summer 2008. I think I can be ready, but the issue is as much about analyzing data as it is about collecting it. Who will I get to help, especially when I am so busy around exam time with administrative duties?
  • What kind of concurrent validity claims might I make? Will I ask teachers to rate their students and then compare performance with expectations, or will I use data that we already collect (classroom scores or rated evaluations)? Are teachers very good at predicting student ability? Lee's 2005 thesis of the L/S tests at the ELC say no (as did major portions of this study), but I think that both of those are problematic: Lee because the teacher ratings were probably based on Speaking when she was comparing those to Listening scores, and in Cumming's case, he admits that BICS/CALP may have a lot to do with false impressions. Of course BICS/CALP may be an issue for me too, but if students and teachers practice items in the classroom more than once, then teachers ought to be good at predicting success, shouldn't they?
  • How will tasks be adjusted for lower levels (adhering the concerns of Sara in this study)? Laura (our Level 1 teacher) has already told me that she thinks Level 1 students need to be able to hear the integrated listening passage twice. As it is, they barely understand it before it's over. A second listening for Levels 1 and 2 might be appropriate and justified (given the findings of this study and Laura's experience).

Tuesday, November 6, 2007

G-Study of iBT

The most recent paper that I finished reading is:

Lee, Y. (2006). Dependability of scores for a new ESL speaking assessment consisting of integrated and independent tasks. Language Testing, 23, 131-166.

Summary: Lee conducts a generalizability study of prototype speaking tasks for the iBT (internet-based TOEFL). The purpose of the study was to assess to what degree reliability increased with the number of tasks and ratings.

Tasks include three types: integrated reading-speaking, integrated listening-speaking, and independent. The effect of rater was ignored due to the difficulties of orchestrating a fully-crossed design; however, rating was considered as a facet, so all ratings were analyzed in a double-rating model.

Lee concludes that increasing tasks has a greater influence on increasing reliability than does increasing raters. Lee also found that there was sufficient evidence to justify collapsing analytic scores into a holistic score of speaking given that correlations among task types were quite high.

This has implications for the iBT speaking component. First, including more than one speaking task is more likely to increase the reliability of scores than adding extra raters. In other words, Lee suggests that more reliable scores are given when raters have a greater number of speaking sample from the examinees than when multiple raters are used. In essence, many raters can still come up with divergent scores when only one sample is used, but when many samples are provided, even a small number of raters is highly accurate at scoring the examinee.

The second implication affects the use of holistic scoring. Rather than require raters to score each sample separately (and then averaging the separate scores into a composite score for each examinee), Lee suggests that since all task types tend to correlate highly, there is sufficient evidence to warrant the use of a holistic score for each examine based on an overall impression rather than a combined average. The benefit is that this reduces the time it takes raters to score a single examinee.

Critique: Admittedly I am not impressed that Lee did not attempt a fully-crossed design. I think this makes the "rater has low effect in increasing reliability" claim given that we have nothing more than a 1-rating versus 2-ratings option. I recognize that fully-crossed designs are really hard to manage, but Lee was funded by ETS (who creates the TOEFL), so I don't know why ETS didn't demand a more rigorous research design.

I am also concerned that Lee does not describe the testing occasions. Did students take all 12 tasks at once, leading to enormous test fatigue? Or were the testing occasions spread out over several weeks thereby possibly confounding test results given that students may have improved in proficiency over the period of testing. I need to email Lee and ask about this specifically.

Connection: Thankfully I have been introduced to generalizability theory in LING 660 (Language Testing) and IP&T 752 (Measurement Theory), because it is hard for a researcher to really explain it well in a short article. I had also previously read a study (for my thesis) written by Rob Schoonen (who is a much better writer than Lee) and this helped me to understand the analysis and results that Lee described.

Lee does a decent job of connecting the results of this study to others. I found myself coming to similar conclusions and making connections to my limited experience with G-studies (such as Schoonen). I also found it valuable to read this study in connection to the iBT validation studies by Cumming. It helped me gain a bigger picture of the iBT validation project and how all these different elements function together to inform the development of this high-stakes exam.

Additional Reflections: This article sparked a lot of ideas for me.
  • Could my study include a comparison of analytic/holistic scoring? Or at the least, could I compare portfolio scores (holistic) with scores for just the integrated tasks? This may justify the need for a separate score for draft writing and timed writing.
  • I like that Lee suggests that test developers need to state the purpose of the task. Are independent tasks just about language, and integrated tasks just about content? In our case, probably not. They are both about language. This need to be clarified.
  • A scoring-related validity component would be a valuable part of my evaluation that might also include teacher and student content/context validity evidence.

Friday, November 2, 2007

Research Progress

I just spoke with Sharon (our programmer) about how the integrated writing tasks are going (we are test piloting them with Levels 1, 3, and 5 today). It was good to see that Level 1 (the only level that handwrote) was able to summarize – and even synthesize – the listening and reading passages. We talked about how Level 3 might be especially difficult (given that it is not Level 3 adjusted and it TOEFL level with just a few vocabulary changes to make it easier).

We also discussed writing ratings and the evaluation study. She is very open to programming an interface for online writing rating feedback since she previously developed something similar to it in the past (for speaking rating). I explained that raters would want to see what others had rated a portfolio after they submitted their own rating to the system. She thought I meant all current raters, but I explained that it would be a database of previous ratings (collected from previous semesters and maybe including a recent set of ratings that I solicit from current teacher to be added to the database of benchmark portfolios).

I also think that the database should include more than just scores; information from the evaluation surveys indicates that raters want to see/hear the justifications for a rating. So, perhaps I could get current raters to write justifications (or maybe audio record them and then transcribe them) and add those to the database along with their scores.

Until this week, the evaluation had been progressing slowly. At the end of last semester I distributed surveys to all raters, but my research assistants never picked up the surveys before they both left for out of town jobs at the end of the semester. Frustrated, but not discouraged, I resent the surveys at the beginning of this semester, but half of the raters had left the ELC (due to graduation and full-time job offers) so I was only able to send out seven surveys. At the beginning of this week I had only received two of the seven, so I sent out another reminder and by this morning I had gotten two more. I will follow-up with the remaining raters in order to get the maximum number of responses possible.

The next step for the evaluation? I need to schedule interviews or a focus group with the raters to follow-up on the interviews. TREC had counseled that I not lead the interviews/focus group in an effort to minimize rater inhibition for free expression. I could a grad student to do it for me, but I want to use someone who is familiar with the rating system. I could also use a current teacher, who is one of my raters, but will her personal bias influence how she interprets feedback from others?

The next step for the dissertation study will involve sending a pilot survey to teachers and students based on the practice LAT is being piloting this week and next. The survey piloting will help me to revise the survey for actual use at the end of the semester during LATs. I have received IRB approval, but I am still awaiting final approval on my release form (the ORCA office wants me to update it).

Friday, October 19, 2007

Super(ior) Talker

Last week I dialed into the Language Testing Institute (LTI) of ACTFL in order to prove that I can, in fact, speak English.

This officially certifiable oral proficiency interview (OPI) by an ACTFL rater is a necessary step to becoming certified to conduct ACTFL OPIs. In addition to speaking hours interviewing and rating ESL speakers, I also need to prove that I meet the ACTFL guidelines for a superior level speaker.

The interview, which didn't make me nervous, but did make me curious due to the fact that I have only ever done this from the interviewer side, started 20 minutes late due to the interviewer's late conducting of the previous interview. But, since I knew what kinds of speech sample the interviewer needed to elicit from me, I think that we were able to make the conversation move along rather quickly, I was surprised to learn that the interview had only taken about 20 minutes (a typical superior level interview can take 30 minutes), but we had covered a variety of topics and had plenty of samples of superior level speech for her to analyze.

The results of the interview were posted today, and I am, in fact, a superior speaker of English. Surprise! I guess all these years of learning and practicing English have really paid off. What a relief.

Of course this is only one small step in this OPI experience. I still need to get the trainer's feedback from my practice round, and then I will have to work at collecting interviews for my certification round (where I actually have to do a good job as opposed to just learning as I did this summer). Here's hoping the trainer doesn't rush to send me the feedback. I've got enough projects to do.

Although I did accomplish to important tasks today: I got the integrated writing practice tasks sent to the programmer (for pilot testing next month) and I submitted my integrate tasks IRB application. It's all working out, despite my poor planning and procrastination.

Saturday, October 13, 2007

Integrated Writing Task

My plans for a dissertation topic revolve around the use of integrated writing tasks. These exam items, used in ESL assessment, require students to write about a reading and listening passage. We are now piloting the use of integrated writing tasks at our center.

Reasons Why We Want to Implement the Integrated Task
  • Every semester we inevitably deal with cases of unintentional plagiarism due to students whose summary, paraphrase, and quoting skills are weak
  • Most of our students are planning to attend university in the USA and need to learn English for Academic Purposes (EAP) which involves writing from sources
  • Integrated writing tasks are now part of the Test of English as a Foreign Language (TOEFL) which many of our students want to take and "pass"
Differences Between the TOEFL's and Our Integrated Tasks
  • Our integrated tasks are used in connection with our Level Achievement Tests in order to assess students' readiness to move on the next class level whereas the TOEFL is used to assess readiness for university study which our program only assesses in an indirect form through level advancement
  • Our implementation of the integrated tasks will assess student ability from novice to low advanced, namely our Levels 1 through 5; the TOEFL is designed to assess a narrower band of language ability, namely the ability of students in our Levels 4 and 5
  • The TOEFL tasks are rated individually, but the integrated tasks will likely become part of our holistic portfolio score, or maybe even a subset portfolio score for fluency (non-draft) writing
In order to properly implement these tasks for our own center, I'm conducting a literature review of relevant studies. The first:
  • Cumming, A., Kantor, R., Baba, K., Erdosy, U., Eouanzoui, K, & James, M. (2005). Differences in written discourse in independent and integrated prototype tasks for next generation TOEFL. Assessing Writing, 10, 5-43.
  • This study analyzed a variety of textual features of prototype integrated tasks versus the current independent writing tasks. It suggests that integrated tasks encourage the use of a greater variety of vocabulary and greater use of source ideas. The independent tasks are better at eliciting more language and at developing better arguments from examinees.
  • I hope to be able to use this study as a model for the quantitative analysis of my dissertation research. Like the Cumming et al. study, I will examine a variety of discourse features and how they differ between responses to our independent and integrated tasks; I will also plan to look at these language features as they differ across our 5 proficiency levels.
I'll add more studies as I read them.

Tuesday, October 9, 2007

Moving forward

I heard back from IRB today. My request to extend my research has been granted. Of course now that I want to do a different (but related) study for my dissertation, I will need to submit a new IRB application.

Why start something new for my dissertation? I'm applying for assistant professor positions (to start next Fall) and one of them is looking for expertise in an area that I haven't focused on very much: oral assessment and listening/speaking pedagogy. Of course I had taught L/S classes, and I'm in the process of becoming OPI certified, but I haven't done much on this topic from a research perspective. So now's my chance.

Truth is, whatever an employer wants me to be an expert in, I can do it. I have only become focused on writing assessment and methodology because:

1) My thesis chair wanted to research it, so I did it for my thesis, and
2) My current job as the writing coordinator requires me to become an expert on it.

So if a potential position requires expertise in something else, I can do it. My revised dissertation project will:

1) Build on my current research,
2) Relate directly to what I do for my job anyway, and
3) Extend my expertise to include research on L/S assessment.

So we'll see how it goes. The fact is, if I want to be ready for these jobs that start Fall 2008, then I need to be all but done my research by then, which gives me less than a year to plan, conduct, and write my dissertation. Heh. Who knows? I could do it. Probably. Maybe. Yeah. And if nothing else, this job application process will motivate me to graduate ahead of schedule. And there's nothing bad about that idea.

Tuesday, September 25, 2007

Working on a List

I'm doing my best to check items off of my To-Do list. Today, I sent a draft of my article off for faculty review, I emailed off my OPI materials, and I renewed my IRB application. I've got a variety of projects that I can work on this semester, but they are coming at the expense of other duties. In addition to projects and teaching, I ought to be doing more to train and support our teachers, but that isn't really happening. At this point I'm mostly expecting teachers to figure it out by working together or reaching out to me proactively. Hopefully I will catch up with my projects and book in time to support teachers.

Friday, September 21, 2007

Another Push in the Write Direction

The hints have no longer come subtlety. In fact they could hardly be called hints. Demands might be too strong, but mandates is probably not far.

In the past three weeks I have had two professors encourage me to get back to work on submitting my thesis research to an academic journal. I told them that I was working on a collaborative article, but the truth is my collaborators are not collaborating, so nothing is happening.

Then I figured it was time to start getting my job application skills back on track since I'll only been in my current position until 2009. And of course that means adding publications to my CV.

And then today my boss strongly, strongly, strongly urged me to get my article published. A political, greater good sort of thing. And so that's what I have spent this evening trying to do. Write.

And it's not really working. But somehow I will manage.

Wednesday, September 19, 2007

OPI Round 1 Deadline

The practice round of my OPI certification is coming up next week. I did the training earlier this summer and I conducted interviews all summer long, but my plan to get everything smoothed out and ready this fall got complicated by my busy teaching schedule.

However, I hope to get it all wrapped up and sent off next week before Kimberly and I head to the Grand Canyon for a short weekend trip.

This may not be the most relaxing semester, but at least we are having lots of learning opportunities.

Tuesday, September 4, 2007

Fresh Start, New Setback

September marks the beginning of a new semester. It will be my first official fall semester as a PhD student after officially being admitted in spring term.

Spring term was useful and busy even though I had the two months off from my job. I spent the time in professional development and in graduate course credits. Summer, however, was more frustrating. I never quite got into a good teaching groove for my teaching, and the graduate course that I took (well, that Kimberly took since the professor held it during a time that I could not meet) was the worst course that I have ever taken. It was poorly organized, poorly articulated, poorly graded, and the instructor still has yet to submit grades even though the term finished about a month ago.

Still, I was ready to let go of the summer and head into a productive fall. Unfortunately I have a heavy teaching load for this semester (my two regular classes, plus a graduate writing class and a a double-section for one of my regular courses), but I am confident that it will work out just fine. I was looking forward to being a student in an interesting technology course, however that plan has been altered. The department, without informing students, has added a required course to the already hefty list of required classes. This new course conflicts with my teaching schedule (yet again), but the instructor has agreed to let me attend for the portions that I can.

I'm not happy with the lack of organization and communication of this department. But then I remind myself that I am not paying for tuition and that regardless of what happens with my credits, it is a great learning opportunity. So long as I approach it with that attitude, I will be less likely to be disappointed and frustrated. And since I never really intended to do my PhD here to begin with, I can feel good about whatever the outcome over the next two years.