nbsmpqmy_wpjaecs2020 Comments on: Plenary Talk 1: ‘Corpus Linguistics’ or ‘Linguistics with a Corpus’? https://jaecs2020.laurenceanthony.net/plenary-talk-1/?utm_source=rss&utm_medium=rss&utm_campaign=plenary-talk-1 Japan Association for English Corpus Studies Annual Conference Sun, 04 Oct 2020 09:16:14 +0000 hourly 1 https://wordpress.org/?v=7.0.2 By: dyskobol https://jaecs2020.laurenceanthony.net/plenary-talk-1/#comment-187 Sun, 04 Oct 2020 06:48:34 +0000 https://www.jaecs2020.org/?p=878#comment-187 Hi,

Many thanks for your answers. As for research material, I meant a type of textual material that you use for your research (corpus, text, text sample etc.) But I agree that a unit of observation basically means the same. I think that character n grams may have linguistic interpretation, for example in morphological research (e.g. when you study morphological productivity), e.g. reuse, where re- stands for ‘again’ etc.). On a different note, there is also psycholinguistic research showing that many word n grams are not holistic units by dint of their high frequency only. That is why I considered your distinction (character n-grams are not meaningful linguistically but word n-grams are) to be arbitrary and hence my question.

I agree that keywords and keyness are abstract concepts that were defined and operationalized differently by researchers (by Williams, Philips, Scott, among others, and by researchers in non-Anglophone countries) and there is no unified method to study them.

Once again, many thanks for your presentation and discussion.

]]>
By: jegbert https://jaecs2020.laurenceanthony.net/plenary-talk-1/#comment-143 Sun, 04 Oct 2020 00:05:27 +0000 https://www.jaecs2020.org/?p=878#comment-143 Hi,

Thank you for watching my talk. Here are some quick responses:

1. ‘Unit of analysis’ is the same as ‘unit of observation’. I use ‘variable’ in its technical statistical/scientific sense. I’m not sure what ‘research material’ refers to.

2. There is extensive research linking word forms and word n-grams to psycholinguistic processes. Word n-grams are functional and they, as units, play a role in language acquisition, use, and processing. Thus, they are linguisically meaningful/interpretable. To my knowledge, no such links have been made with character n-grams. I’m not saying character n-grams could not have a linguistic interpretation, only that I’ve never heard one and it’s difficult to imagine what that might be.

3. Keyness is an abstract construct, not a specific corpus method. It was first proposed by Firth and further developed by Williams many decades ago. They used the term to refer to the ‘aboutness’ of a text or discourse domain. ‘Corpus frequency keyness’ was the first corpus-based operationalization of the construct of keyness, proposed by Mike Scott in the early 1990s. That was one way to operationally realize the construct of keyness, but certainly not the only way. It is simply so deeply embedded in the research culture of corpus linguistics that some have come to think there is no other way keyness could possibly be operationalized. This is simply untrue. In fact, a few years later Scott proposed a competing measure: key keyword analysis. Our method is an additional operational method that attempts to capture keyness (or aboutness). These measures can be compared side by side to determine which is most appropriate and effective for a particular research goal.

]]>
By: dyskobol https://jaecs2020.laurenceanthony.net/plenary-talk-1/#comment-138 Sat, 03 Oct 2020 19:50:10 +0000 https://www.jaecs2020.org/?p=878#comment-138 That is a very interesting talk, thank you. I have a few questions. First, do you think that labels such as ‘unit of analysis’ and ‘research material’ would be more intuitive rather than ‘variable’ and ‘unit of observation’? I find the latter one quite confusing. Second, as for validity, could you elaborate more on why you consider character n grams less linguistically meaningful than word n grams (cf. th vs the of the). What may count here is whether a unit of analysis/variable is a form-meaning mapping, I guess, which is relative to research goal and scope). And third, as for keyness, isn’t it a problem of how you define keywords or keyness? If we accept a definition whereby a keyword is a word that occurs in one corpus with outstanding frequency as compared to another corpus, then it is difficult to avoid using frequency as a determining factor. If you uses text dispersion (or range), isn’t it the case that you then search for something else (be it a content-distinctive words or otherwise). Anyway, you have raised a number of important issues and your talk has given me a lot of food for thought. Thank you once again.

]]>