摘要
The need within language testing and assessment for specifications of test content, methods and constructs has long been recognised and accepted. In education more generally, statements of aims, objectives and curricular frameworks are also widely provided, although the definitions and operationalisations of these may vary greatly. In second and foreign language education, there is a long tradition of stating expected levels of achievement, with or without reference to curricular objectives. However, such stated levels were often vague and were not defined independently of higher or lower levels. Hence, it was very common to use terms like Beginner, False Beginner, High or Low Intermediate and Advanced or Mastery, the meanings of which were interpreted very differently in different contexts, cultures and educational systems, such that one person’s or system’s “False Beginner” might be another’s “Intermediate”. More recently, efforts have been made to define achievement levels in terms of the requirements of educational contexts or of learners’ needs and their ability to function independently in particular settings (see, for example, the Threshold level developed by language specialists under the aegis of the Council of Europe in the 1970s or the Vantage and Mastery levels in the 1980s). Draft versions of the Common European Framework of Reference were developed in the 1990s, which speculated on or stated what learners at particular levels could do with the particular language being assessed. The DIALANG suite of diagnostic language tests in 14 European languages ( https://dialangweb.lancaster.ac.uk/ ) was based on a draft (1996) version of the Common European Framework of Reference for Languages (CEFR) and used Can Do statements taken from the CEFR as its basis for test item development and the reporting of test results. Partly as a result of the success and take-up of DIALANG, the CEFR was rapidly adopted in many European educational systems, and CEFR-based exit levels from secondary or higher education were required by law in some countries (e.g. the Austrian Ministry of Education and the University of Innsbruck as reported in Spöttl et al. 2016 ). This requirement to use the CEFR was not without controversy (see Fulcher, 2004 , for example, or Weir, 2005 ), but their arguments were largely ignored by language educators. Language policies were widely developed with the CEFR at their core, both within and beyond Europe, resulting not only in CEFR-based language tests, but also CEFR-based language textbooks and language education curricula. Despite, or perhaps because of, the wide and rapid adoption of the CEFR, the somewhat inevitable limitations of the CEFR became apparent, and these are described in the paper by Jin, Wu, Alderson and Song in this collection of papers. As a result of these perceived limitations, Japan developed its own localised version of the CEFR, the CEFR-J (see Jin et al.). More recently, the Chinese Government, aware of the existence, nature and impact and spread of the CEFR, decided to develop its own national framework and standards for English. The National Education Examinations Authority (NEEA) was charged with developing a theoretical model for a national framework (NEEA, 2014), which quickly became known as the China’s Standards of English (CSE). Questions that readers may ask themselves when reading this Special Issue include How does the CEFR influence the CSE, or perhaps, How does the CSE draw on and use/diverge from/improve the CEFR? What evidence (empirical and theoretical) for validity is presented or needed in the CSE? What data is available on the reliability and levels of descriptors? What information is available in the various papers on the actual content of scales, test types, test features, test development and design? What are the issues in localisation and prescription of the CSE across China? How are the tests to be developed based on the CSE, and how will they be implemented and administered in practice? What problems emerge and how are they addressed? What impact and washback will CSE-related tests have? The paper by Jin et al. is a key document that will repay careful attention and re-reading, as it makes the case for a national framework for English education in China. It spells out the need for such a framework, the potential benefits of the China’s Standards of English, the challenges of such a project and potential pitfalls. It provides a critique of the Common European Framework of Reference (CEFR) and argues for a national rather than an international framework. It acknowledges that such an ambitious project will encounter resistance and criticisms from vested interests and foresees that macro-politics as well as micro-politics will play an important role in this. The paper is structured around four key questions: Why do we need a national framework of reference for English in China? Why do we need a national framework of reference for English in China? Why do we create a new framework instead of adopting or adapting an existing one? Why do we create a new framework instead of adopting or adapting an existing one? What are the challenges facing government organisations in the process of constructing and implementing a national framework? What are the challenges facing government organisations in the process of constructing and implementing a national framework? What are the challenges facing the individuals involved in or having stakes in the construction and development of a national framework? What are the challenges facing the individuals involved in or having stakes in the construction and development of a national framework? It is hoped that a national framework will provide a shared understanding of the standards of English education and assessment, that it will also have a positive impact on the quality of teaching, learning and assessment of English and that it “will better prepare Chinese people to live and work in an increasingly globalised world”. One may legitimately wonder whether the CSE project is re-inventing the wheel and whether, after all the effort and resources going into it have been expended, the resulting CSE will be significantly different from the CEFR. Only time will tell, but this is a brave attempt to innovate in what will certainly involve difficult circumstances in such a huge and varied country. Differences between the Chinese and the European contexts are acknowledged, and some of the limitations of the CEFR are discussed and will hopefully be avoided. The authors of this lead paper remain positive but acknowledge that the project will need to continue to develop, be revised and followed up with a range of accompanying developmental and research projects. The paper by He and Chen describes progress in the development of a set of scales for Listening in a second or foreign language. This will certainly have to be referred to in any similar future project not only in listening but also in other language skills, including Reading of course, but also Grammar and Mediation at the very least. The basic constructs investigated are cognitive ability, listening strategies, linguistic knowledge and performance in typical listening activities. Cognitive ability is to be assessed by scales of narration, description, exposition, argumentation, instruction and interaction. Listening strategy is addressed by scales of planning, execution and evaluation and repair. Typical listening activities will involve scales of listening to conversations, listening to lectures, listening to announcements and instructions, listening to broadcasts and watching movies and TV series. Use of linguistic knowledge will involve scales of grammar and pragmatics. Interestingly, vocabulary is not addressed in this paper on listening but in a separate paper in the collection by Zhao, Wang, Coniam and Xie. This article reports a descriptive approach to scale development, which presented a number of difficulties that have yet to be resolved. The four research questions that were formulated were: How do we define the construct of listening ability with respect to the English teaching and learning context of China? How do we define the construct of listening ability with respect to the English teaching and learning context of China? How do we describe listening ability in a comprehensive manner? How do we describe listening ability in a comprehensive manner? How do we collect descriptors with reasonable representativeness? How do we collect descriptors with reasonable representativeness? How do we scale the descriptors and how to validate the scales? How do we scale the descriptors and how to validate the scales? The data collection itself was extensive. The researchers examined 42 documents in both English and Chinese, including proficiency scales, teaching syllabuses, curriculum requirements, test specifications and rating scales, resulting in 1240 descriptors. They invited 159 teachers, 475 students and 119 professionals from different fields to write descriptors, which resulted in 1263 descriptors. Nationwide large-scale questionnaire surveys are planned to be distributed to 10,000 teachers and 100,000 students in over 1000 schools. I look forward to reading the results of this ongoing research in due course. The paper by Zeng and Fan on developing scales for Reading is in many ways parallel to the paper on Listening, since they both deal with language comprehension, but in other ways, it is different, especially in the research questions asked. The authors give a very brief account of the growing number of proficiency scales around the world, while pointing out that China does not have nationwide proficiency scales. They introduce the CEFR which is being drawn upon for the development of the CSE but argue that there is a need for localisation of the CEFR which would better reflect the current practice of learning and teaching of English in China as well as having positive effects on English learning and teaching. Two questions are addressed in the paper, namely “What is the theoretical basis for developing reading scales?” and “What are the parameters for describing reading proficiency?” Core members of the project designed the methodology and trained teachers to compile and develop reading descriptors. They were also responsible for collecting relevant literature and analysing interviews with experts and the descriptors developed by a working group consisting of 12 teacher leaders, 94 other teachers and a team of experts in language testing or second language education. Useful details are given of the documents analysed (which are largely similar to those that the original developers of the CEFR also consulted, with the addition of some Chinese documents) and the procedures developed for drafting descriptors and for refining the parameters for describing reading proficiency. Figure 1 gives a useful overview of this process. Similar to the Listening descriptor development team, the Reading project members developed a framework for the CSE Reading scales consisting of Cognitive Ability, Comprehension Strategy and Knowledge, and some details are given of the components of these constructs. Similar to the development of the Listening descriptors, a large number (14,467) of reading descriptors was gathered, of which 1398 were summarised by team members from the literature and 13,069 were compiled by teachers. Again, useful details are given of the process of refining the parameters for classification of the descriptors and for revising and removing those draft descriptors that did not conform to these parameters. This process initially resulted in 4884 descriptors which were further reduced to a total of 574 descriptors to be entered in the database. Finally, the paper briefly discusses the improvement of the theoretical basis of the resulting scales, the removal of references to native-speaker standards and norms for reading proficiency (which also happened in the development of the CEFR scales) and the grouping of the descriptors according to the function intended by each descriptor. This paper gives a fairly detailed outline of the theoretical basis for the procedures, constructs and parameters for developing reading descriptors, without, however, much practical exemplification of the results. Such exemplification will eventually have to be achieved in future research and development by showing how and why descriptors were accepted, rejected or revised during the process outlined in this paper. This research should involve both qualitative and quantitative empirical studies, interviews and questionnaire surveys and analyses and standard-setting studies as mentioned below in order to identify proficiency levels and cut scores. Such research will certainly need to extend well beyond the currently envisaged deadline of 2017. Unlike other papers in this Special Issue, the article by Liu and Jia does not look at the development of descriptors, scales and test constructs but rather seeks to validate a localised university-based speaking test which is based on the CEFR. In contrast to the paper by He and Chen, this is a small study involving 54 learners, two interlocutors and two raters. Moreover, the researchers analyse transcriptions of videoed performances of only one part of the speaking assessment, with first year university students. Unusually, the researchers concentrated on counting the language functions displayed by the students and explored which features of speaking distinguish speaking proficiency at each level. In the end, however, it proved impossible to relate functions and features like length of turns, choice of words and syntax, hesitation markers and topic coherence to speakers’ level of oral proficiency. This is nevertheless a potentially interesting task, which might offer lessons on test development relevant to other researchers concerned with the CSE. It is often asserted that teaching and assessing Chinese learners of English neglects the speaking skill in order to concentrate on the testing of reading, writing and grammar. To the extent that this is true, then efforts to validate the assessment of the ability to speak English are laudable. As Luoma ( 2004 : 1) points out, however, “Assessing speaking is challenging…because there are so many factors that influence our impression of how well someone can speak a language, and because we expect test scores to be accurate, just and appropriate for our purpose. This is a tall order.” This paper illustrates some of the difficulties involved in assessing speaking. attempt to speaking a understanding of the nature of as described in and of speaking in a second or foreign language, which are in the constructs and that that It also developing speaking that to how and why people speak in any language. speaking the development of test or assessment need to be developed to or the assessment of a person’s speaking also need to how performance on achievement and proficiency tests of speaking can be reported to learners and their teachers, and research needs to how well the test results learners’ performance in the are to to especially on a and speaking The paper by and The authors on a case study to whether and to what extent the CEFR is or can be to be the construction of a writing ability scale for English in Chinese The two research questions have two as with descriptors from other to what extent are CEFR descriptors in describing English writing What can English do in CEFR writing descriptors for English with descriptors from other to what extent are CEFR descriptors in describing English writing What can English do in What can English do in CEFR writing descriptors for English CEFR writing descriptors for English How can we construct a writing ability scale of Can Do statements for English by descriptors from various What are the procedures to be followed in developing a Can Do descriptor What should be adopted to scale the descriptors? How can we construct a writing ability scale of Can Do statements for English by descriptors from various What are the procedures to be followed in developing a Can Do descriptor What are the procedures to be followed in developing a Can Do descriptor What should be adopted to scale the descriptors? What should be adopted to scale the descriptors? In one could argue that these are not research since potential are available in the CEFR The is whether such procedures can and results in the Chinese In other can the procedures used in developing the CEFR be in the Chinese The paper provides a of a approach to these issues in four 1 is the drafting of a questionnaire Can Do statements taken from the CEFR, to the of English and the as well as the results of group as in in One was the of the questionnaire to teachers and students and the of an descriptor involved teachers the level of of descriptors with reference to writing at levels of taken from four of In interviews were with university teachers for their the levels of the descriptors in order to the of writing ability scale what students can do at different levels of One the and the results of the Two the of the writing and the original questionnaire as This detailed of the results is for its not only in the of the of the but also and perhaps more in of this study in other and Moreover, the detailed of this methodology will similar studies at lower levels of English education in China. The version of the writing presented in 1 , of with descriptors at the level descriptors at and 14 descriptors at descriptors the ability level required for students of English, that of students of English and 1 are what in English are to be As with the other papers in this Special Issue, it to be whether further empirical evidence or this into levels and whether the intended levels are of the actual writing performance of English in China. It is that the of English education across such a varied as China might both the and the results. this paper is a and basis for in future The paper by Zhao, Wang, Coniam and the as a reference and the of the study as an of the CSE vocabulary descriptors with reference to the CEFR vocabulary They acknowledge at the however, that some of the CEFR descriptors have been for for in describing vocabulary for of of terms and for the which that it has to the nature of vocabulary in particular or the nature of 2005 : Moreover, the authors acknowledge that of the CEFR and the English of ) that the vocabulary descriptors in both documents were to describe vocabulary knowledge for English education in these it is perhaps to this paper as a but rather as an interesting study which some of the CEFR vocabulary descriptors to whether Chinese teachers of English at level can on the levels at which the descriptors set in the are for vocabulary in the CSE. In the results of the study the teachers as to the levels of the descriptors and or to One of the for such might be that the were with the CEFR the or that the particular descriptors for the study were of the descriptors for the study proved to range over more than levels were one level higher or one level lower than their original CEFR ). that this was a study to the of the CEFR and vocabulary descriptors to develop the CSE scales, the authors that of the scales provides evidence of the CSE scales, and can be a of information for the future improvement of the However, the authors the methodology by not only quantitative and analyses of the but also to in in to their They also more and the CEFR descriptors, and a of the CEFR descriptors it is somewhat that a study will provide evidence of The is that is not for these and that of evidence should also be to the empirical evidence so that and can be This is a very ambitious with a of only from to This the actual of the large and empirical to have been in and the of the first version of the CSE in 2017. In this Special Issue of is a on a that will for many not to up various on the data and to the first set of scales and develop To give a the CEFR was in after two draft versions been in the were both that and are ongoing include work on new descriptors for and as well as of new for developing such frameworks and also of the original The in the paper on English vocabulary education to all of the empirical studies in this Special Issue, or envisaged for the development of the CSE. of testing and will be needed to develop more comprehensive descriptors which are to reflect the range of CSE scales. will then need to be to empirical and further of One that is mentioned in this yet to is that of levels. the of language proficiency and how it in learners over the effort to distinguish different levels of language proficiency is not without One important is what does it to be a and how is a to be It is interesting to that the paper on listening separate levels of listening ability the number of levels in the the paper on speaking test only levels of speaking ability and the paper on the writing scales only levels of writing this to be a then it needs to be for example, that the 54 students in the paper on speaking were students in English, and they a with the large range of ability to speak English across the Chinese education It is that a large number of levels or of language proficiency will be needed to this how levels can be it is that it in one of the effects of the CEFR that much research has at how to set standards which one might include and how to develop cut scores. DIALANG 2005 ) was of a in and implementing standard-setting procedures in the in order to whether a given is at or the levels according to the CEFR and the scales, Can Do statements and proficiency descriptors. in and analysing standard-setting procedures and the of cut scores has for example, the research by ( , ), the by and ( ) and ( ) and the of The European of Language and on standard-setting on its and especially the Special on the CEFR ( ). I that the CSE would from further of this This is an important I look forward to reading future reports of empirical research which will to its