Coming up in BYU-COCA, mini-interview with Prof. Mark Davies

I had the great pleasure to be able to put 4 questions to Professor Mark Davies about the BYU corpora tools and the upcoming changes (

1. Would you share what you think are interesting site user statistics?

There’s some good data at; let me know if you have questions about any of that.

2. What is the motivation behind the upcoming user interface changes?

The main thing is to have an interface that works well on laptops/desktops, as well as tablets, as well as mobile devices (cell phones, etc). More and more people are connecting to websites that work well as they are on-the-go with mobile devices, and the current corpus interface doesn’t work well for that.

3. Any chance of screenshot previews of the upcoming user interface?

Please see attached. As you can see, each of the four main “pages” — search, results, KWIC, and help — are in their own “page”, which takes up the whole screen. Users click on the tabs at the top of the page to move between these (just as they would click on the different “frames” in the existing interface). But because each “page” takes up the entire page, it will still work fine with cell phones, for example (where frames don’t work well at all).

The other cool thing in the new interface is the ability to create and use “virtual corpora” (see for their implementation in Wikipedia corpus, using the current interface).


4. One tweeter (@cainesap) was wondering in what register would cat pictures go in the web genre corpus?

🙂 🙂 Good question. See the core.png file attached 🙂


The first screenshot indicates the new interface looks much cleaner and easier to use from mobiles devices and as a bonus my ebook – Quick Cups of COCA won’t be needing too much change for the new edition : )

Thanks to Professor Davies for taking time to answer this mini-interview and to you for reading.

Bonus Questions!

5. Do the search results screens stay the same as now?

Pretty similar; yes.

6. Still a mystery where cat pics and graphical memes, animated gifs go? : )​

This and many other equally profound questions can be answered with the new corpus :-).

Corpus Linguistics for Grammar – Christian Jones & Daniel Waller interview

CLgrammarFollowing on from James Thomas’s Discovering English with SketchEngine and Ivor Timmis’s Corpus Linguistics for ELT: Research & Practice I am delighted to add an interview with Christan Jones and Daniel Waller authors of Corpus Linguistics for Grammar: A guide for research.

An added bonus are the open access articles listed at the end of the interview. I am very grateful to Christian () and Daniel for taking time to answer my questions.

1. Can you relate some of your background(s)?

We’ve both been involved in ELT for over twenty years and we both worked as teachers and trainers abroad for around a decade; Chris in Japan, Thailand and the UK and Daniel in Turkey. We are now both senior lecturers at the University of Central Lancashire (UCLan, Preston, UK),  where we’ve been involved in a number of programmes including MA and BA TESOL as well as EAP courses.

We both supervise research students and undertake research. Chris’s research is in the areas of spoken language, corpus-informed language teaching and lexis while Daniel focuses on written language, language testing (and the use of corpora in this area) and discourse. We’ve published a number of research papers in these areas and have listed some of these below. We’ve indicated which ones are open-access.

2. The focus in your book is on grammar could you give us a quick (or not so quick) description of how you define grammar in your book?

We could start by saying what grammar isn’t. It isn’t a set of prescriptive rules or the opinion of a self-appointed expert, which is what the popular press tend to bang on about when they consider grammar! Such approaches are inadequate in the definition of grammar and are frequently contradictory and unhelpful (we discuss some of these shortcomings in the book).  Grammar is defined in our book as being (a) descriptive rather than prescriptive (b) the analysis of form and function (c) linked at different levels (d) different in spoken and written contexts (e) a system which operates in contexts to make meaning (f) difficult to separate from vocabulary (g) open to choice.

The use of corpora has revolutionised the ways in which we are now able to explore language and grammar and provides opportunities to explore different modes of text (spoken or written) and different types of text. Any description of grammar must take these into account and part of what we wanted to do was to give readers the tools to carry out their own research into language. When someone is looking at a corpus of a particular type of text, they need to keep in mind the communicative purpose of the text and how the grammar is used to achieve this.

For example, a written text might have a number of complex sentences containing both main and subordinate clauses. It may do so in order to develop an argument but it can also be more complex because the expectation is that a reader has time to process the text, even though it is dense, unlike in spoken language. If we look at a corpus we can discover if there is a general tendency to use a particular pattern such as complex sentences across a number of texts and how it functions within these texts.

3. What corpora do you use in the book?

We have only used open-access corpora in the book including BYU-BNC, COCA, GloWbe, the Hong Kong Corpus of Spoken English. The reason for using open-access corpora was to enable readers to carry out their own examinations of grammar. We really want the book to be a tool for research.

4. Do you have any opinions on the public availability of corpora and whether wider access is something to push for?

Short answer: yes. Longer answer: We would say it’s essential for the development of good language teaching courses, materials and assessments as well as democratising the area of language research. To be fair to many of the big corpora, some like the BNC have allowed limited access for a long time.

5. The book is aimed at research so what can Language Teachers get out of it?

By using the book teachers can undertake small-scale investigations into a piece of language they are about to teach even if it is as simple as finding out which of two forms is the more frequent. We’ve all had situations in our teaching where we’ve come across a particular piece of language and wondered if a form is as frequent as it is made to appear in a text-book, or had a student come up and say ‘can I say X in this text’ and struggled with the answer. Corpora can help us with such questions. We hope the book might make teachers think again about what grammar is and what it is for.

For example, when we consider three forms of marry (marry, marries and married) we find that married is the most common form in both the BYU-BNC newspaper corpus and the COCA spoken corpus. But in the written corpus, the most common pattern is in non-defining relative clauses (Mark, who is married with two children, has been working for two years…). In the spoken corpus, the most common pattern is going to get married e.g. When are they going to get married?

We think that this shows that separating vocabulary and grammar is not always helpful because if a word is presented without its common grammatical patterns then students are left trying to fit the word into a structure and in fact words are patterned in particular ways. In the case of teachers, there is no reason why an initially small piece of research couldn’t become larger and ultimately a publication, so we hope the book will inspire teachers to become interested in investigating language.

6. Anything else you would like to add?

One of the things that got us interested in writing the book was the need for a book pitched at undergraduate students in their final year of their programme and those starting an MA, CELTA or DELTA programme who may not have had much exposure to corpus linguistics previously. We wanted to provide tools and examples to help these readers carry out their own investigations.

Sample Publications

Jones, C., & Waller, D. (2015). Corpus Linguistics for Grammar: A guide for Research. London: Routledge.

Jones, C. (2015).  In defence of teaching and acquiring formulaic sequences. ELT Journal, 69 (3), pp 319-322.

Golebiewksa, P., & Jones, C. (2014). The Teaching and Learning of Lexical Chunks: A Comparison of Observe Hypothesise Experiment and Presentation Practice Production. Journal of Linguistics and Language Teaching, 5 (1), pp.99–115. OPEN ACCESS

Jones, C., & Carter, R. (2014). Teaching spoken discourse markers explicitly: A comparison of III and PPP. International Journal of English Studies, 14 (1), pp.37–54. OPEN ACCESS

Jones, C., & Halenko, N.(2014). What makes a successful spoken request? Using corpus tools to analyse learner language in a UK EAP context. Journal of Applied Language Studies, 8(2), pp. 23–41. OPEN ACCESS

Jones, C., & Horak, T. (2014). Leave it out! The use of soap operas as models of spoken discourse in the ELT classroom. The Journal of Language Teaching and Learning, 4(1), pp.1–14. OPEN ACCESS

Jones, C, Waller, D., & Golebiewska, P. (2013). Defining successful spoken language at B2 Level: Findings from a corpus of learner test data. European Journal of Applied Linguistics and TEFL, 2(2), pp.29–45.

Waller, D., & Jones, C. (2012). Equipping TESOL trainees to teach through discourse. UCLan Journal of Pedagogic Research, 3, pp. 5–11. OPEN ACCESS

Discovering English with SketchEngine – James Thomas interview

2015 seems to be turning into a good year for corpus linguistics books on teaching and learning, you may have read about Ivor Timmis’s Corpus Linguistics for ELT: Research & Practice. There is also a book by Christian Jones and Daniel Waller called Corpus Linguistics for Grammar: A guide for research.

This post is an interview with James Thomas,, on Discovering English with SketchEngine.

1. Can you tell us a bit about you background?

2. Who is your audience for the book?

3. Can your book be used without Sketch Engine?

4. How do you envision people using your book?

5. Do you recommend any other similar books?

6. Anything else you would like to add?

1. Can you tell us a bit about your background?^

Currently I’m head of teacher training in the Department of English and American Studies, Faculty of Arts, Masaryk University, Czech Republic. In addition to standard teacher training courses, I am active in e-learning, corpus work and ICT for ELT. In 2010 my co-author and I were awarded the ELTon for innovation in ELT publishing for our book, Global Issues in ELT. I am secretary of the Corpora SIG of EUROCALL, and a committee member of the biennial conference, TALC (Teaching and Language Corpora).

My work investigates the potential for applying language acquisition and contemporary linguistic findings to the pedagogical use of corpora, and training future teachers to include corpus findings in their lesson preparation and directly with students.

In 1990, I moved to the Czech Republic for a one year contract with ILC/IH and have been here ever since. Up until that time, I had worked as a pianist and music teacher, and had two music theory books published in the early 1990s. Their titles also beginning with “Discovering”! 🙂

2. Who is your audience for the book?^

The book uses the acronym DESKE. Quite a broad catchment area:

  • Teachers of English as a foreign language.
  • Teacher trainees – the digital natives – whether they are doing degree courses or CELTA TESOL Trinity courses.
  • People doing any guise of applied linguistics that involve corpora.
  • Translators, especially those translating into their foreign language. (Only yesterday I presented the book at LEXICOM in Telč.)
  • Students and aficionados of linguistics.
  • Test writers.
  • Advanced students of English who want to become independent learners.

3. Can your book be used without Sketch Engine?^

No. (the answer to the next question explains why not).

Like any book it can be read cover to cover, or aspects of language and linguistics can be found via the indices: (1) Index of names and notions, (2) Lexical focus index.

4. How do you envision people using your book?^

It is pretty essential that the reader has Sketch Engine open most of the time. Apart from some discussions of features of linguistic and English, the book primarily consists of 342 language questions/tasks which are followed by instructions – how to derive the data from the corpus recommended for the specific task, and then how to use Sketch Engine tools to process the data, so that the answer is clear.

Example questions:
About words
Can you say handsome woman in English?
Do marriages break up or down?
How is friend used as a verb?
Which two syllable adjectives form their comparatives with more?
Do men say sorry more than women?

About collocation
I’ve come across boldly go a few times and wonder if it is more than a collocation.
It would be reasonable to expect the words that follow the adverb positively
to be positive, would it not?
Is there anything systematic about the uses of little and small?
What are some adjectives suitable for giving feedback to students?

About phrases and chunks
Does at all reinforce both positive and negative things?
What are those phrase with lastleast; believeears; leadhorse?
How do the structures of to photograph differ from take a photo(graph),
guess with make a guess, smile with give a smile?
Which –ing forms follow verbs like like?

About grammar
How do sentences start with Given?
Who or whom?
Which adverbs are used with the present perfect continuous?
Do the subject and verb typically change places in indirect questions?
How new and how frequent is the question tag, innit?

About text
Are both though and although used to start sentences? Equally?
How much information typically appears in brackets?
Does English permit numbers at the beginning of sentences?
Is it really true that academic prose prefers the passive?
In Pride and Prejudice, are the Darcies ever referred to with their first names?

There is an accompanying website with a glossary – a work eternally in progress, and a page with all the links which appear in the footnotes (142 of them), and another page with the list of questions, which a user might copy and paste into their own document so that they can make notes under them.

5. Do you recommend any other similar books?^

The 223 page book has three interwoven training goals, the upper level being SKE’s interface and tools, the second being a mix of language and linguistics, while the third is training in deriving answers to pre-set questions from data.

AFAIK, there is nothing like this.

6. Anything else you would like to add?^

In all the conference presentations and papers and articles that I have seen and heard over the years in connection with using corpora in ELT, with very few exceptions teachers and researchers focus on a very narrow range of language questions. When my own teacher trainees use corpora to discover features of English in the ways of DESKE, they realise that the steep learning curve is worth it. They are being equipped with a skill for life. It is a professional’s tool.

Sketch Engine consists of both data and software. Both are being constantly updated, which argues well for print-on-demand. It’ll be much easier to bring out updated versions of DESKE than through standard commercial publishers. I’m also expecting feedback from readers, which can also be incorporated into new editions.

My interests in self-publishing are partly related to my interest in ICT. This book is printed through the print-on-demand service, One of the beauties of such a mode of publishing is the relative ease with which the book can be updated as the incremental changes in the software go online. This is in sharp contrast to the economies of scale that dictate large print runs to commercial publishers and the standard five-year interval between editions.

There is a new free student-friendly interface which has its own corpus and interface, known as SKELL which has been available for less than a year. It is also undergoing development at the moment, and I will be preparing a book of worksheets for learners and their teachers (or the other way round). I see it as a 21st cent. replacement of the much missed “COBUILD Corpus Sampler”.

Lastly, I must express my gratitude to Adam Kilgarriff, who owned Sketch Engine until his death from cancer on May 16th, at the age of 55. He was a brilliant linguist, teacher and presenter. He bought 250 copies of my book over a year before it was finished, which freed me up from other obligations – a typical gesture of a wonderful man, greatly missed.

Many thanks to James for taking the time to be interviewed but pity my poor wallet with some very neat CL books to purchase this year. James also mentioned that, for a second edition file, Chapter 1 will be re-written to be able to use the open corpora in SketchEngine.

Skylight interview with Gill Francis & Andy Dickinson

Skylight is a relatively new corpus interface designed with teachers and students in mind. Gill Francis one of the developers kindly answered some questions. The news about forthcoming suggestions for classroom activities is something to look forward to as well as the collocation feature. It is interesting to note that Gill is very much in favour of the use of keyword in context (KWIC) concordance lines. Others such as the FLAX language learning team see KWICs as more of an hinderance and propose their own novel interfaces.

Can you share a little of your background?

Andrew Dickinson is a software writer who is interested in the use of corpora in the classroom and Gill Francis (that’s me) is a corpus linguist. In 1991 I joined the pioneering Cobuild project as Senior Grammarian. Cobuild was founded in 1980 by Professor John Sinclair (University of Birmingham). Its aim was to compile and investigate huge collections of written and spoken language in order to produce a range of dictionaries and grammars for learners that reflect how English is actually spoken and written today. My interest and direction in corpus linguistics owes everything to John Sinclair and our colleagues at Cobuild.

The Bank of English corpora grew to about 450 million words by the late 1990s. We used a fast, versatile, and powerful corpus analysis tool called ‘lookup’. As a grammarian, I was responsible for the grammatical information in the second edition of the Collins Cobuild Advanced Learner’s Dictionary (1995), along with Susan Hunston and Elizabeth Manning. The three of us also wrote the Cobuild Grammar Patterns series (1996, 97, and 98). All these publications reflected a detailed study of corpus evidence.

I’ve continued to work and publish in corpus linguistics since leaving Cobuild. (A list of publications is available.) Then a few years ago I got together with Andy to design Skylight, a program with a clear, easy interface for use by teachers and learners. Since then we have presented Skylight at various corpus linguistics conferences and seminars, and are currently developing it for more general release.

You are targeting classroom use by teachers with Skylight so what do you hope to bring that other corpus tools don’t?

1 – A clear, simple interface

Skylight has a clear, visually attractive interface. The query language is simple and intuitive, and can be learned in a couple of minutes. You can make a query by simply typing in a word or phrase without any special spacing or punctuation, for example “in my opinion” or “in the middle of” or “it’s a case of”.

To vary any word in the query, you use a pipe: “in my|his|her opinion”, or “in the middle|midst of”.

If you want to vary the query and see the range of words in a particular phrase or frame, you use one or more asterisks, for example “in my * opinion” will return “in my humble opinion”, “in my honest opinion”, “in my personal opinion” and so on.

This is about as complex as the query language gets – click on the User Manual from any page of Skylight to see examples of each kind of query. The rules are few and easily mastered by teachers and learners.

2 – Fast, easy alphabetical sorting

If you want to sort concordance lines to the right, or the left, you just click on a button above the lines. This helps you to see at a glance what the right-hand or left-hand collocates of a word or phrase are.

3 – Worksheets and classroom activities

If you are a teacher, you can use Skylight to prepare your own worksheets for corpus-based language activities. When you receive the results of a query, you can tailor the lines to fit your teaching point. This means that you can show only the lines you want, or hide those that you don’t, by clicking or entering text. You can copy the result into Word or another application using the Copy to Clipboard button. The results appear as a neat table, properly displayed and ready for your use. See the User Manual for further details and lots of examples.

Ideally, too, teachers and learners would be able to access a corpus at any point during a class, whenever they want to investigate how a word or phrase is used in a range of real language texts and situations.

For initial guidance and ideas, we are also preparing a large number of suggestions for stand-alone classroom activities practising points of grammar, lexis, and phraseology. Some of these activities address language change and the tension between prescription and description in language teaching. We’ll let you know when we release the first batch of these.

4 – A range of corpora

There are several corpora already available on Skylight – choose any one from the drop-down menu. For example, there is a very large general corpus, ukWaC, which contains 1.4 billion words, as well as smaller corpora like the BNC, BASE, and VOICE. Then there are even smaller corpora – for example a corpus of all Shakespeare’s plays and sonnets that is particularly useful for school children studying English literature.

In addition, any corpus can be compiled in response to the needs of groups of users, such as English school children or intermediate level EFL students. This depends, of course, on copyright restrictions. For more information, see the final sections of the User Manual.

Which other corpus tools would you recommend for teachers either in the classroom or outside?

We don’t feel particularly qualified to answer this question. There a lot of tools that access huge corpora and are extremely useful to linguists and lexicographers, such as Sketch Engine; the COCA (a large corpus of American English) concordancer, and Lancaster’s Corpus Query Processor. If you look up ‘corpus’ and ‘classroom’ together in any search engine, there will be several hits, but we don’t know of anything that combines an easy-to-use interface with really good classroom applications. This doesn’t mean there isn’t anything of course!

What present and/or future do you see for Google as a corpus in language learning?

One of the drawbacks of compiled corpora, such as UkWaC and the BNC, is that they are a snapshot of how language is used at a particular time (or at successive times, if a corpus is updated on a regular basis). The gathering and cleaning-up of text can take many months, so all corpora – even the most recent – are necessarily out-of-date by the time they appear.

The only way to get today’s language today is to use the web as a corpus (see for example Birmingham City University’s WebCorp). This gives results in the KWIC (Key Word in Context) format, with the word or phrase in the centre. The results are not cleaned up or processed, however, which limits their usefulness in the classroom.

But Google itself won’t give you the output you need for focusing on a word or phrase, sorting it, or looking at collocations. You’ll get plenty of examples, of course, but they won’t be shown in the KWIC format. The KWIC display is probably the most important and exciting development in modern corpus linguistics, and you need it if you are to do real corpus-based language work in the classroom or anywhere else.

Anything else you would like to add?

You asked whether we intend to add information about collocation. We are experimenting with a display modelled on the ‘Picture’ technique used in the lookup software used for the Bank Of English, which shows where collocates appear in relation to the node (the central word or phrase) – whether they tend to occur before or after it, for example.

We call the collocation display ‘Searchlight’. The Searchlight display below shows that the most frequent words immediately after obvious are that, then reasons (plural), then choice, then reason (singular). The most frequent words two to the right are of, for, and is. And so on – the columns are not connected, of course; they simply give positional collocations.

The brilliant thing about ‘picture’ that we want to replicate is that you simply click on any word to go to the relevant concordance lines. So if you click on reasons, you’d get all the lines with the combination obvious reasons. So it gives you a subset of the lines, which can then be sorted and tailored in any way you like.


We will add Searchlight to the Skylight website as soon as possible, though we have not yet decided whether to add statistical information – probably not. In the meantime, I’d just like to say that in my many years of scrolling down concordance lines, I find that alphabetical sorting is a very good guide to the collocations of a word. I happened to search for the word intuitively recently, and returned 500 lines. If I sort them one to the right and scroll rapidly down, it’s clear that among the most frequent adjectives that follow it are appealing, correct, and obvious, while the verbs are know and understand. If I sort them one to the left, it is clear that one of the most frequent collocates is the verb be in various forms: ‘it is intuitively obvious’ and so on. Sorting one way and the other gives you a quick thumbnail sketch of a word, and is extremely useful.

So go ahead and try Skylight. And above all, click onto the User Manual, which tells you all you need to know and provides lots of examples of searches using different features.

A huge thanks to the Skylight team and do comment here about your opinions of the interface.

Thanks for reading.