<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="4.4.1">Jekyll</generator><link href="https://nina-nullspace.netlify.app/feed.xml" rel="self" type="application/atom+xml" /><link href="https://nina-nullspace.netlify.app/" rel="alternate" type="text/html" /><updated>2026-08-05T15:16:13+00:00</updated><id>https://nina-nullspace.netlify.app/feed.xml</id><title type="html">The Nina Null Space</title><subtitle>AI/NLP researcher, accessibility practitioner, and musician based in Copenhagen.</subtitle><entry><title type="html">Four Steps Ahead - Advice For My Past Self on Structuring a PHD</title><link href="https://nina-nullspace.netlify.app/productivity/2022/01/31/Four-Steps-Ahead.html" rel="alternate" type="text/html" title="Four Steps Ahead - Advice For My Past Self on Structuring a PHD" /><published>2022-01-31T19:44:00+00:00</published><updated>2022-01-31T19:44:00+00:00</updated><id>https://nina-nullspace.netlify.app/productivity/2022/01/31/Four-Steps-Ahead</id><content type="html" xml:base="https://nina-nullspace.netlify.app/productivity/2022/01/31/Four-Steps-Ahead.html"><![CDATA[<h2 id="introduction">Introduction</h2>

<p>As of this writing, I have been a PHD student for almost 5 months, and I’ve learned more about myself and how I operate in terms of work routines and structures than I ever did as a university student.
I have had a mild burn-out already from over-working myself, and I have also experienced not feeling like I had any energy going forward (we’ll talk about perfectionism in a later post, not to worry).</p>

<p>In the following, though, I’ll be outlining the four biggest insights I’ve had as a PHD student in my last 5 months with respect to structuring my day.</p>

<h2 id="treat-your-phd-as-a-job">Treat your PHD as a job</h2>

<p>A PHD-program is a full-time job. Yes, you have a lot of freedom and a lot of options for how you want to structure a PHD, particularly if you are working on a solo research project through a stipend from the university like I am, as opposed to a research group or an externally funded project with potential stakeholders and conflicts of interest that might narrow your options and hold you more accountable to a certain extend. Yes, it is also an education, and the point of taking an education is to learn and be more process-oriented, at least ideally.
But no, it is a job, and you have to approach it as such. You’ll be faced with a variety of different tasks and cooperations you have to manage, and you need to get good at structuring your day in accordance with what works for you. You’ll have external obligations and numerous of people to be accountable to, so your work day needs to intersect with that of other people\s at least to a certain extend, even if you choose to work from home.</p>

<p>One particular challenge I am still working to solve is to find a way to divide how much time I spend on different activities in accordance with my temperament. I have never been very good at multi-tasking; I prefer to work on one thing at a time in a very linear fashion, but that just isn’t feasible when you have 3-4 different activities to manage each week.
As an example, the following is my general tasks in the next 6 months, and the allocated amount of hours (This is of course a rough outline, but this presupposes that I have a 40 hour work week and need to structure my tasks in accordance with this limitation.):</p>

<ul>
  <li>teaching a beginner programming course once a week, along with preparing the material; approximately 9 hours/week</li>
  <li>Working on a technical course project to gain the necessary skills (and ECTS credits) I need; approximately 8 hours/week</li>
  <li>Tweaking experiments and preparing an article about evaluation of a lexical sentiment resource to publish at an international journal, approximately 8 hours/week</li>
  <li>Gathering data for my research project in hyperbole detection; approximately 8 hours/week</li>
  <li>Reading theoretical articles about my subject of choice; approximately 4 hours/week</li>
  <li>Administrative work, including meetings, responding and following up on emails, and applications for funding for external research visit; approximately 3 hours/week</li>
</ul>

<p>Have I followed my own advice so far? No, not as much as I could have. I have fallen into the trap multiple times of prioritizing things that interest me the most and piling up the mundane things, and as far as work/life-balance goes, let’s just say that in the first half of january, I worked 12 hours/7 days a week for at least 2 weeks in a row to complete a deadline, and then crashed hard for the remainder. I get very hyper-focused on whatever I’m focusing on, for better or worse.</p>

<h2 id="get-some-sleep">Get some sleep!</h2>

<p>Seriously, this one is crucial. You might think that it doesn’t matter if you stay up until 3 at night to work on a project with a deadline, but if there is something I’m starting to notice after implementing better sleeping habits, it is that consistency is important.
A recent <a href="https://pubmed.ncbi.nlm.nih.gov/31583118/">study</a>, finding strong correlations between sleep quality, consistency, and duration and study performance.
One major point that I found interesting is that sleeping well the night before a significant milestone, like an academic test, did not correlate with performance. However, consistently good sleep appeared to massively affect performance.</p>

<p>I realize there could be a certain ‘duh’ factor in the statement that sleep is very, very important, but as we just saw, this is a genuine problem for many students and adults alike. My own history with sleep has been buggy; it’s as though my brain lights up at night when it has sometime to relax and wander off into the distance. Sometimes it wanders to speculating on a research problem, like the role of formal semantics in NLP and how to utilize predicate logical structures in claim strength detection, or it jumps to the far future and speculates on the use of bio-weapons and gene editing technologies, or it speculates on patterns and insights in my private life. In the worst case, it starts an elaborate analysis of all the things I still need to do, and beats down any solutions with excessively perfectionistic standards, or it will invent problems to solve that prooves I am not at all doing enough with my life.
In short, my brain can go anywhere, and it really, really wants to party.</p>

<p>I’m finding that taking sometime off in the hours before going to sleep to relax, read a good book, or meditate/reflect substantially reduces my mental monolog. I’m also implementing better phone habits, as I do tend to want to check my news feed or talk to people, which actually seems to influence my ability to sleep, particularly if a lot of new information and data is coming in.</p>

<p>Another thing that can help with reducing mental chatter is exercise, which is why I’ve implemented a 20 minute routine before I sleep to tire myself out. Yoga stretches, weight-lifting, or general agility exercises seem to work to remind me that I have a body that needs things, and sometimes it needs sleep. Again, ‘duh’, you say, and now go sit in the corner and be quiet while I finnish this post.</p>

<h2 id="realize-that-youre-going-to-make-plans-and-they-are-going-to-change">Realize that you’re going to make plans, and they are going to change</h2>

<p>When I first applied for my PHD, I had to make a project proposal and a rough plan for each semester, detailing key goals and milestones on my project, courses I wanted to teach/take, and where I wanted to go for the compulsory guest research stay.
Regardless of what kind of program you are taking, I’ll venture to say that planning is going to come in at some very early point in the process.
Awesome, right? You made a plan, and now you can stick to it and make sure you build your perfect schedule, because you can predict the future and know exactly what you want to be doing 5 years ahead, but you didn’t consider all the little bumps along the <em>and</em> you forgot to eat for the last 3 days. Well, that is, if you are me.
You might very well also think that making plans is completely unnecessary to start with. As usual, the truth lies somewhere in the middle: Plans ensure that you have a rough idea of how your program is going to look, and at the very least, it shows the people assessing your potential as a PHD candidate that you are able to realistically plan your schedule in accordance with the allocated time.
However, once you start the program, don’t expect that plan to remain stationary. You’ll run into scheduling conflicts, unanswered questions that you need people to help you with, and just straight-up things that are not possible within the time frame.</p>

<p>One of the most impactful advice in project management that I’ve seen is to realize you’re failing early, then readjust.
I’ve had to readjust many plans already, and I’ve also gotten stuck in processing decisions way longer than I ought to have, even if my perception of the situation that lead me to the decision has usually not been wrong. Usually, this is because of one of two things:</p>

<ol>
  <li>I don’t have enough information to go through with the decision, and I’m procrastinating on getting it because it requires me to explain myself in a phone call or an email.</li>
  <li>There legitimately is no better decision than the other, and I just have a hard time figuring out which one will get me to where I want to be.</li>
</ol>

<p>The advice I would give my past self, even as of one week ago, and which I was also given by somebody very close to me, is to actively stop yourself from processing while you are in problem-solving mode rather than rumination mode, and start doing. Get the information you need, or make a final decision and don’t look back.</p>

<h2 id="finally-do-whatever-works-for-you">Finally, do whatever works for you</h2>

<p>In this post, I’ve outlined some of the things I’ve had to massively adjust since starting my PHD, and I’ve written this with the intention of giving useful pointers. I also find that writing them down as advice provides me with a detached perspective on what works for me. So, if you find this post far too one-sided or you know these things (A) won’t work for you, or (B) are not a problem for you whatsoever, great!
Society has all kinds of metrics with which to evaluate work productivity and work ethic. However, when it comes down to it, the most important thing is to find your own flow and do what you know works for you, not what any productivity books or PHD moles tell you works for you. Be experimental, figure things out as you go, and you’ll be fine, but be honest about whether what you are doing is truly fascilitating your goals. What do you really want?
In the end, that is really only a question you yourself can answer.</p>]]></content><author><name></name></author><category term="Productivity" /><summary type="html"><![CDATA[Introduction As of this writing, I have been a PHD student for almost 5 months, and I’ve learned more about myself and how I operate in terms of work routines and structures than I ever did as a university student. I have had a mild burn-out already from over-working myself, and I have also experienced not feeling like I had any energy going forward (we’ll talk about perfectionism in a later post, not to worry). In the following, though, I’ll be outlining the four biggest insights I’ve had as a PHD student in my last 5 months with respect to structuring my day. Treat your PHD as a job A PHD-program is a full-time job. Yes, you have a lot of freedom and a lot of options for how you want to structure a PHD, particularly if you are working on a solo research project through a stipend from the university like I am, as opposed to a research group or an externally funded project with potential stakeholders and conflicts of interest that might narrow your options and hold you more accountable to a certain extend. Yes, it is also an education, and the point of taking an education is to learn and be more process-oriented, at least ideally. But no, it is a job, and you have to approach it as such. You’ll be faced with a variety of different tasks and cooperations you have to manage, and you need to get good at structuring your day in accordance with what works for you. You’ll have external obligations and numerous of people to be accountable to, so your work day needs to intersect with that of other people\s at least to a certain extend, even if you choose to work from home. One particular challenge I am still working to solve is to find a way to divide how much time I spend on different activities in accordance with my temperament. I have never been very good at multi-tasking; I prefer to work on one thing at a time in a very linear fashion, but that just isn’t feasible when you have 3-4 different activities to manage each week. As an example, the following is my general tasks in the next 6 months, and the allocated amount of hours (This is of course a rough outline, but this presupposes that I have a 40 hour work week and need to structure my tasks in accordance with this limitation.): teaching a beginner programming course once a week, along with preparing the material; approximately 9 hours/week Working on a technical course project to gain the necessary skills (and ECTS credits) I need; approximately 8 hours/week Tweaking experiments and preparing an article about evaluation of a lexical sentiment resource to publish at an international journal, approximately 8 hours/week Gathering data for my research project in hyperbole detection; approximately 8 hours/week Reading theoretical articles about my subject of choice; approximately 4 hours/week Administrative work, including meetings, responding and following up on emails, and applications for funding for external research visit; approximately 3 hours/week Have I followed my own advice so far? No, not as much as I could have. I have fallen into the trap multiple times of prioritizing things that interest me the most and piling up the mundane things, and as far as work/life-balance goes, let’s just say that in the first half of january, I worked 12 hours/7 days a week for at least 2 weeks in a row to complete a deadline, and then crashed hard for the remainder. I get very hyper-focused on whatever I’m focusing on, for better or worse. Get some sleep! Seriously, this one is crucial. You might think that it doesn’t matter if you stay up until 3 at night to work on a project with a deadline, but if there is something I’m starting to notice after implementing better sleeping habits, it is that consistency is important. A recent study, finding strong correlations between sleep quality, consistency, and duration and study performance. One major point that I found interesting is that sleeping well the night before a significant milestone, like an academic test, did not correlate with performance. However, consistently good sleep appeared to massively affect performance. I realize there could be a certain ‘duh’ factor in the statement that sleep is very, very important, but as we just saw, this is a genuine problem for many students and adults alike. My own history with sleep has been buggy; it’s as though my brain lights up at night when it has sometime to relax and wander off into the distance. Sometimes it wanders to speculating on a research problem, like the role of formal semantics in NLP and how to utilize predicate logical structures in claim strength detection, or it jumps to the far future and speculates on the use of bio-weapons and gene editing technologies, or it speculates on patterns and insights in my private life. In the worst case, it starts an elaborate analysis of all the things I still need to do, and beats down any solutions with excessively perfectionistic standards, or it will invent problems to solve that prooves I am not at all doing enough with my life. In short, my brain can go anywhere, and it really, really wants to party. I’m finding that taking sometime off in the hours before going to sleep to relax, read a good book, or meditate/reflect substantially reduces my mental monolog. I’m also implementing better phone habits, as I do tend to want to check my news feed or talk to people, which actually seems to influence my ability to sleep, particularly if a lot of new information and data is coming in. Another thing that can help with reducing mental chatter is exercise, which is why I’ve implemented a 20 minute routine before I sleep to tire myself out. Yoga stretches, weight-lifting, or general agility exercises seem to work to remind me that I have a body that needs things, and sometimes it needs sleep. Again, ‘duh’, you say, and now go sit in the corner and be quiet while I finnish this post. Realize that you’re going to make plans, and they are going to change When I first applied for my PHD, I had to make a project proposal and a rough plan for each semester, detailing key goals and milestones on my project, courses I wanted to teach/take, and where I wanted to go for the compulsory guest research stay. Regardless of what kind of program you are taking, I’ll venture to say that planning is going to come in at some very early point in the process. Awesome, right? You made a plan, and now you can stick to it and make sure you build your perfect schedule, because you can predict the future and know exactly what you want to be doing 5 years ahead, but you didn’t consider all the little bumps along the and you forgot to eat for the last 3 days. Well, that is, if you are me. You might very well also think that making plans is completely unnecessary to start with. As usual, the truth lies somewhere in the middle: Plans ensure that you have a rough idea of how your program is going to look, and at the very least, it shows the people assessing your potential as a PHD candidate that you are able to realistically plan your schedule in accordance with the allocated time. However, once you start the program, don’t expect that plan to remain stationary. You’ll run into scheduling conflicts, unanswered questions that you need people to help you with, and just straight-up things that are not possible within the time frame. One of the most impactful advice in project management that I’ve seen is to realize you’re failing early, then readjust. I’ve had to readjust many plans already, and I’ve also gotten stuck in processing decisions way longer than I ought to have, even if my perception of the situation that lead me to the decision has usually not been wrong. Usually, this is because of one of two things: I don’t have enough information to go through with the decision, and I’m procrastinating on getting it because it requires me to explain myself in a phone call or an email. There legitimately is no better decision than the other, and I just have a hard time figuring out which one will get me to where I want to be. The advice I would give my past self, even as of one week ago, and which I was also given by somebody very close to me, is to actively stop yourself from processing while you are in problem-solving mode rather than rumination mode, and start doing. Get the information you need, or make a final decision and don’t look back. Finally, do whatever works for you In this post, I’ve outlined some of the things I’ve had to massively adjust since starting my PHD, and I’ve written this with the intention of giving useful pointers. I also find that writing them down as advice provides me with a detached perspective on what works for me. So, if you find this post far too one-sided or you know these things (A) won’t work for you, or (B) are not a problem for you whatsoever, great! Society has all kinds of metrics with which to evaluate work productivity and work ethic. However, when it comes down to it, the most important thing is to find your own flow and do what you know works for you, not what any productivity books or PHD moles tell you works for you. Be experimental, figure things out as you go, and you’ll be fine, but be honest about whether what you are doing is truly fascilitating your goals. What do you really want? In the end, that is really only a question you yourself can answer.]]></summary></entry><entry><title type="html">A Glance at the Uniform Information Density Hypothesis in Language Processing and Speech Production</title><link href="https://nina-nullspace.netlify.app/emnlp-2021/2022/01/30/A-Glance-at-the-Uniform-Information-Dencity-Hypothesis.html" rel="alternate" type="text/html" title="A Glance at the Uniform Information Density Hypothesis in Language Processing and Speech Production" /><published>2022-01-30T19:30:34+00:00</published><updated>2022-01-30T19:30:34+00:00</updated><id>https://nina-nullspace.netlify.app/emnlp-2021/2022/01/30/A-Glance-at-the-Uniform-Information-Dencity-Hypothesis</id><content type="html" xml:base="https://nina-nullspace.netlify.app/emnlp-2021/2022/01/30/A-Glance-at-the-Uniform-Information-Dencity-Hypothesis.html"><![CDATA[<h2 id="introduction">Introduction</h2>

<p>If you are reading this post, I will, until proven otherwise, assume that you are human, and that you are belonging to the subset of humans who are very interested in natural language processing. As a human, when you are comprehending these sentences, we can reasonably say that you are engaging in a complex act of communication that is, in some way, shaped by some universal set of processes that underlie human cognition. As a formally trained linguist that branched out into AI through my master’s program, I’m fascinated with a subset of natural language processing that uses computational models and vast amounts of data to look for universal human characteristics when it comes to how we process language. This field, formally known as computational psycholinguistics, deals with questions such as the following:</p>

<ul>
  <li>Is there a connection between processing effort when reading a text, and how unexpected a word/sentence being read is?</li>
  <li>Is there a connection between the length of production of particular sounds, or phones, and how frequently they appear in the phonetic inventory of a language?</li>
  <li>How universal are such tendencies across languages?</li>
</ul>

<p>This line of research is, sadly, under-represented in AI research communication, but I have found two papers introduced in the Linguistic Theory section of the <a href="https://2021.emnlp.org/papers">2021 Conference in Empirical Methods in Natural Language processing (EMNLP)</a> which I will summarize here, each dealing with exactly those questions with a slightly different emphasis: The first one, <a href="https://aclanthology.org/2021.emnlp-main.74/">“Revisiting the Uniform Information Density Hypothesis”</a>, deals with human text processing, and the other, <a href="https://aclanthology.org/2021.emnlp-main.73/">“A surprisal–duration trade-off across and within the world’s languages”</a>, concentrates on human speech production. Both of those papers center around the very interesting but slightly esoteric-sounding information theoretic concept known as the Uniform Information Density Hypothesis (UID). The UID hypothesizes that language users, on some level of language structure, prefer information to be distributed evenly across a linguistic signal. If that sounds vague or dense, just stick with me and it will all make sense soon enough. If you’re already familiar with the UID, feel free to skip the next section of this post, in which I’ll explain the intuitions behind UID and how it relates to language processing. After that, I will detail the experiments in each of the research papers and their results and make some brief concluding remarks.</p>

<h2 id="a-gentle-overview-of-the-uniform-information-density-hypothesis-uid">A Gentle Overview of the Uniform Information Density Hypothesis (UID)</h2>

<p>As promised, let’s first dig into the origins and assumptions postulated by the Uniform Information Density Hypothesis. The origins of UID are to be found within information theory, a field that had its beginnings in mathematics and communications and has subsequently been applied in multiple disciplines spanning from engineering to cognitive science. The underlying assumption in information theory, first proposed by  <a href="https://en.wikipedia.org/wiki/Claude_Shannon">Claude Shanon</a> (1948), is that communication can be understood as information transmitted at a certain rate limited by some upper-bound of a metaphorical noisy communication channel. With respect to language, what this means is that there exists some maximum capacity through which information content can be communicated reliably while minimizing the error across a linguistic signal. Thus, according to most interpretations, what UID claims is that since there is an upper bound on information rate that humans can reliably handle, they must modulate their flow of information such that it maximizes such a channel capacity, i.e. information density doesn’t ever stray too far from some global mean value of information rate, avoiding an overflow or under-use of the channel at any given point.</p>

<p>So, how is this abstract notion of information density actually measured? Through a concept known as surprisal, which is simply just defined as the negative log probability of some linguistic unit conditioned on its prior context. What this means is that, given some kind of linguistic signal, which for now can be anything from a sentence to the sequence of speech sounds; the higher the probability of some individual linguistic unit occurring in that context, e.g. a word or a sound, the lower the surprisal, i.e. the informational content, becomes. This idea that high-surprisal items carries more informational content reflects the linguistic intuition that, for instance, unpredictable words carry more information than predictable ones in a given sentence, or, as we’ll see the second paper, people spend longer time on producing less frequent speech sounds.</p>

<p>Okay, I did say a <em>gentle</em> overview, so how about an analogy: Imagine the flow of information coming at you in this post is a river. Most of the words and concepts are easy to comprehend, which are nice smooth spots on our river of information. Inevitably, you will encounter words or concepts that are more difficult to comprehend, which are rocks and rapids in the river; these are the surprisals. The UID states that, to make things as easy as possible to float down this river of information, the surprisals shouldn’t all be loaded in one part of this blog post like a giant impassable waterfall, but should be spread out so we can make it through the rapids (even if there is still some difficulty!).</p>

<p>Now, another very relevant question that may arise is, how do you estimate surprisal, given that there is no way to measure the ground-truth probability of a unit in context? This is where natural language processing finally comes in! Surprisal can simply be estimated using the output of a computational language model. Traditionally, N-gram models have been used, although recently, as we will see in the first paper, large pre-trained language models based on transformer architectures have been shown to have supperior psychometric predictive power. In speech data, a sequence-to-sequence architecture such as LSTM is more useful, as further seen in the section on the second paper.</p>

<p>The take-away here is that it’s possible, through the use of computational models and corpora that contains various kinds of psychometric data, i.e. reading times or running speech, to more elaborately verify the intuitions of UID in both language production and language processing.</p>

<p>In the remainder of this post, I’ll introduce each of the two aforementioned papers, which deal with various effects of UID on language comprehention and production, and explain their main empirical contributions and findings.</p>

<h2 id="revisiting-the-uniform-information-density-hypothesis">Revisiting the Uniform Information Density Hypothesis</h2>

<p>Let’s first explore a paper that deals with investigating the implications of UID on language comprehention and linguistic acceptability, a somewhat less explored area than implications on language production. Specifically, the authors explore the relationship between sentence-level processing effort and informational content; i.e. the assummption that the more unexpected a word is in a sentence, the heavier the cognitive load for the reader. They also dig into perceived linguistic acceptability, which in linguistics is simply the extend to which a sentence is permissible by the users of a language as defined by rules of grammaticality. They propose that a sentence that is deemed more acceptable is also easier to process, and vice versa.</p>

<p>To make this more concrete, they provide two example sentences that illustrate how uniform information density might manifest:</p>

<ol>
  <li>“How big is the family that you cook for?”</li>
  <li>“How big is the family you cook for?”</li>
</ol>

<p>Intuitively, they argue, most English-speakers would prefer the 1st variant of the sentence where the relative clause marker, “that”, is included. The information theoretical explanation for this is simply that if “that” is removed, “you” then carries both the marking of a relative clause, as well as the signal of 2nd person singular/plural. Hence, including the relative clause marker spreads information more evenly across a sentence and avoids rapid switching between dense and less dense information.</p>

<p>For data, the authors use multi-modal reading times (self-paced reading times and eye movements) from 4 different corpora, as well as human judgements of linguistic acceptability from another two corpora for English and Dutch. The motivation behind their experiments is to fit different mathematical functions of surprisal on reading time and acceptability datasets to see which one best predicts the psychometric data. In very basic terms, verifying UID would imply that this mathematical function is super-linear: For the reading time data, they ask whether the processing effort of a linguistic signal is a function of the sum of individual surprisals, which would imply that the processing effort increases linearly with the informational content. However, this seems very counter-intuitive, because it would imply that the distribution of information across a sentence does not matter for the processing at all, which multiple psycho-linguistic experiments suggest isn’t the case. If the function is instead super-linear, the high-surprisal utterances would require a disproportionately high processing effort, which would motivate the smoothing information across a linguistic signal and, as such, confirm the UID hypothesis. The same logic then holds for the case of linguistic acceptability.</p>

<p>They explore these different hypotheses by fitting different regression models on the data. To estimate surprisal, they use three pre-trained neural language models; namely BERT, GPT-2, and TransformerXL. The predictive power is then given by the log probabilities under each of these regression models, namely linear and logistic regression, with or without the sum-of-surprisals term. The results are, for the most part, consistent with UID. However, the super-linear effects are far more consistently observed in linguistic acceptability data, suggesting that uniform information density is much more strongly correlated with linguistic acceptability than with reading times. Thus, this might suggest that when judging the grammaticality of a sentence, language users have a stronger preference for uniform distributions of information, although these results are merely correlational and don’t make strong causal claims.</p>

<p>This brings up another interesting question, namely how to determine the scope of uniformity across a given signal: Assuming that UID implies some kind of regression towards a mean information rate, does this apply more globally, within a language in general, or more locally, within a particular sentence? If a global interpretation of UID is more predictive of the given data, then uniformity is a kind of smoothing effect to a global mean rate that information must not heavily deviate from. However, if a more local interpretation holds, uniformity should be perceived less as a smoothing effect, and more as a pressure to avoid shifting rapidly between content of different information densities, as exemplified by the two sentences earlier in this section. The authors investigate this, among others, by exploring the effect of varying the context window size on the change of variability of the computed information rate, and conclude strongly that the global (smoothing) interpretation wins over the local one. This suggests that the initial explanation of UID in this post as maximizing the use of a metaphorical noisy channel, is, indeed, the most fitting one. As such, language users have a preference for smoothing over the entire language, rather than a on a sentence or phrase-level.</p>

<h2 id="a-surprisal-duration-trade-off-across-and-within-the-worlds-languages">A Surprisal-Duration Trade-Off Across and Within the World’s Languages</h2>

<p>So far we’ve delved into text processing and explored some interesting findings about reading time and linguistic acceptability data. We have also established that linguistic users have a preference for smoothing information across the entire language, as opposed to on a phrase, sentence, or document-level, confirming the most popular interpretation of UID as maximizing the use of a metaphorical noisy communication channel. The second paper summarized in this post investigates UID from a slightly different angle, speech production, and addresses an important question about the linguistic universality of UID interpretations by extending the findings to a corpus of 600 different languages. Basically, this paper explores the relationship between informational content and speech production rate, and defines information content, i.e. surprisal, as the negative log probability of one individual sound, a so-called phone, given a sequence of speech sounds.</p>

<p>The hypothesis is that, assuming that the channel capacity of human information processing is roughly the same across languages, it can be expected that languages place various cognitive constraints on how humans smooth the signal over the phonetic inventory of a language in speech production. Namely, the authors demonstrate strong evidence for a trade-off between the surprisal of a phone and the time it takes to produce it; in other words, speakers slow down significantly when producing highly surprising speech sounds, and speed up on more predictable ones. They demonstrate that this condition holds within 319 languages, in which more surprising phones are pronounced with longer duration. Furthermore, they demonstrate that word-initial phones, i.e. sounds occurring at the beginning of a word, are on average more surprising.</p>

<p>For data, the authors use a phone-aligned corpus containing readings of the Bible in 600 languages spread over 70 language families. The data contains 20 hours of spoken word on average per language, making their findings the “most representative evidence for UID to date”. The surprisal estimates were computed by a phone-level LSTM model, and the phone durations were given by automatically generated alignments in the running text.</p>

<p>The most interesting finding of this paper is that this condition holds not only within one given language, but across languages. Understanding what this means requires a bit of understanding about language universals. Within the world’s languages, there are some tendencies with respect to the features shared by many, although not all, languages; for instance, with respect to phonetics, most languages contain syllables constructed of vowels and consonants where the vowel is the nucleus of the syllable, and most languages contain nasal consonants, [m], [n], or [ng]. Thus, their finding implies that languages containing more universally surprising phones compensate by lengthening the duration of their utterances. Furthermore, the authors didn’t find a single instance in which the oppoosite effect was true; i.e. where shorter phone durations were linked to higher surprisal, and they demonstrate through controlling for several potential confounding variables that this analysis holds remarkably well across languages.</p>

<h2 id="concluding-remarks">Concluding remarks</h2>
<p>Each of the summarized papers provide evidence of varying strength for the existance of UID and, as such, make very important contributions to computational psycholinguistics. An area of research that might be interesting to expand upon is how well the findings for language comprehention and linguistic acceptability holds across non-Indo-European languages, in order to avoid Anglocentric biases. However, this might be challenging because these languages are generally under-represented when it comes to resources and language models.</p>

<p>I hope that through this post, I have given you a better understanding of computational psycholinguistics and how natural language processing is used to further research into the nature of linguistic data. I hope that I have concvinced you that this is not only interesting to a particular subset of academics dealing with linguistics and cognitive science, which one might assume given that this topic does not have immediate industrial applications, but also has interest more generally in the NLP space.
Finally, I hope that the journey along this admittedly somewhat dense river of information has been mostly gentle, smooth sailing, and that you at least enjoyed the rock maneuvering along the way.</p>]]></content><author><name></name></author><category term="EMNLP-2021" /><summary type="html"><![CDATA[Introduction If you are reading this post, I will, until proven otherwise, assume that you are human, and that you are belonging to the subset of humans who are very interested in natural language processing. As a human, when you are comprehending these sentences, we can reasonably say that you are engaging in a complex act of communication that is, in some way, shaped by some universal set of processes that underlie human cognition. As a formally trained linguist that branched out into AI through my master’s program, I’m fascinated with a subset of natural language processing that uses computational models and vast amounts of data to look for universal human characteristics when it comes to how we process language. This field, formally known as computational psycholinguistics, deals with questions such as the following: Is there a connection between processing effort when reading a text, and how unexpected a word/sentence being read is? Is there a connection between the length of production of particular sounds, or phones, and how frequently they appear in the phonetic inventory of a language? How universal are such tendencies across languages? This line of research is, sadly, under-represented in AI research communication, but I have found two papers introduced in the Linguistic Theory section of the 2021 Conference in Empirical Methods in Natural Language processing (EMNLP) which I will summarize here, each dealing with exactly those questions with a slightly different emphasis: The first one, “Revisiting the Uniform Information Density Hypothesis”, deals with human text processing, and the other, “A surprisal–duration trade-off across and within the world’s languages”, concentrates on human speech production. Both of those papers center around the very interesting but slightly esoteric-sounding information theoretic concept known as the Uniform Information Density Hypothesis (UID). The UID hypothesizes that language users, on some level of language structure, prefer information to be distributed evenly across a linguistic signal. If that sounds vague or dense, just stick with me and it will all make sense soon enough. If you’re already familiar with the UID, feel free to skip the next section of this post, in which I’ll explain the intuitions behind UID and how it relates to language processing. After that, I will detail the experiments in each of the research papers and their results and make some brief concluding remarks. A Gentle Overview of the Uniform Information Density Hypothesis (UID) As promised, let’s first dig into the origins and assumptions postulated by the Uniform Information Density Hypothesis. The origins of UID are to be found within information theory, a field that had its beginnings in mathematics and communications and has subsequently been applied in multiple disciplines spanning from engineering to cognitive science. The underlying assumption in information theory, first proposed by Claude Shanon (1948), is that communication can be understood as information transmitted at a certain rate limited by some upper-bound of a metaphorical noisy communication channel. With respect to language, what this means is that there exists some maximum capacity through which information content can be communicated reliably while minimizing the error across a linguistic signal. Thus, according to most interpretations, what UID claims is that since there is an upper bound on information rate that humans can reliably handle, they must modulate their flow of information such that it maximizes such a channel capacity, i.e. information density doesn’t ever stray too far from some global mean value of information rate, avoiding an overflow or under-use of the channel at any given point. So, how is this abstract notion of information density actually measured? Through a concept known as surprisal, which is simply just defined as the negative log probability of some linguistic unit conditioned on its prior context. What this means is that, given some kind of linguistic signal, which for now can be anything from a sentence to the sequence of speech sounds; the higher the probability of some individual linguistic unit occurring in that context, e.g. a word or a sound, the lower the surprisal, i.e. the informational content, becomes. This idea that high-surprisal items carries more informational content reflects the linguistic intuition that, for instance, unpredictable words carry more information than predictable ones in a given sentence, or, as we’ll see the second paper, people spend longer time on producing less frequent speech sounds. Okay, I did say a gentle overview, so how about an analogy: Imagine the flow of information coming at you in this post is a river. Most of the words and concepts are easy to comprehend, which are nice smooth spots on our river of information. Inevitably, you will encounter words or concepts that are more difficult to comprehend, which are rocks and rapids in the river; these are the surprisals. The UID states that, to make things as easy as possible to float down this river of information, the surprisals shouldn’t all be loaded in one part of this blog post like a giant impassable waterfall, but should be spread out so we can make it through the rapids (even if there is still some difficulty!). Now, another very relevant question that may arise is, how do you estimate surprisal, given that there is no way to measure the ground-truth probability of a unit in context? This is where natural language processing finally comes in! Surprisal can simply be estimated using the output of a computational language model. Traditionally, N-gram models have been used, although recently, as we will see in the first paper, large pre-trained language models based on transformer architectures have been shown to have supperior psychometric predictive power. In speech data, a sequence-to-sequence architecture such as LSTM is more useful, as further seen in the section on the second paper. The take-away here is that it’s possible, through the use of computational models and corpora that contains various kinds of psychometric data, i.e. reading times or running speech, to more elaborately verify the intuitions of UID in both language production and language processing. In the remainder of this post, I’ll introduce each of the two aforementioned papers, which deal with various effects of UID on language comprehention and production, and explain their main empirical contributions and findings. Revisiting the Uniform Information Density Hypothesis Let’s first explore a paper that deals with investigating the implications of UID on language comprehention and linguistic acceptability, a somewhat less explored area than implications on language production. Specifically, the authors explore the relationship between sentence-level processing effort and informational content; i.e. the assummption that the more unexpected a word is in a sentence, the heavier the cognitive load for the reader. They also dig into perceived linguistic acceptability, which in linguistics is simply the extend to which a sentence is permissible by the users of a language as defined by rules of grammaticality. They propose that a sentence that is deemed more acceptable is also easier to process, and vice versa. To make this more concrete, they provide two example sentences that illustrate how uniform information density might manifest: “How big is the family that you cook for?” “How big is the family you cook for?” Intuitively, they argue, most English-speakers would prefer the 1st variant of the sentence where the relative clause marker, “that”, is included. The information theoretical explanation for this is simply that if “that” is removed, “you” then carries both the marking of a relative clause, as well as the signal of 2nd person singular/plural. Hence, including the relative clause marker spreads information more evenly across a sentence and avoids rapid switching between dense and less dense information. For data, the authors use multi-modal reading times (self-paced reading times and eye movements) from 4 different corpora, as well as human judgements of linguistic acceptability from another two corpora for English and Dutch. The motivation behind their experiments is to fit different mathematical functions of surprisal on reading time and acceptability datasets to see which one best predicts the psychometric data. In very basic terms, verifying UID would imply that this mathematical function is super-linear: For the reading time data, they ask whether the processing effort of a linguistic signal is a function of the sum of individual surprisals, which would imply that the processing effort increases linearly with the informational content. However, this seems very counter-intuitive, because it would imply that the distribution of information across a sentence does not matter for the processing at all, which multiple psycho-linguistic experiments suggest isn’t the case. If the function is instead super-linear, the high-surprisal utterances would require a disproportionately high processing effort, which would motivate the smoothing information across a linguistic signal and, as such, confirm the UID hypothesis. The same logic then holds for the case of linguistic acceptability. They explore these different hypotheses by fitting different regression models on the data. To estimate surprisal, they use three pre-trained neural language models; namely BERT, GPT-2, and TransformerXL. The predictive power is then given by the log probabilities under each of these regression models, namely linear and logistic regression, with or without the sum-of-surprisals term. The results are, for the most part, consistent with UID. However, the super-linear effects are far more consistently observed in linguistic acceptability data, suggesting that uniform information density is much more strongly correlated with linguistic acceptability than with reading times. Thus, this might suggest that when judging the grammaticality of a sentence, language users have a stronger preference for uniform distributions of information, although these results are merely correlational and don’t make strong causal claims. This brings up another interesting question, namely how to determine the scope of uniformity across a given signal: Assuming that UID implies some kind of regression towards a mean information rate, does this apply more globally, within a language in general, or more locally, within a particular sentence? If a global interpretation of UID is more predictive of the given data, then uniformity is a kind of smoothing effect to a global mean rate that information must not heavily deviate from. However, if a more local interpretation holds, uniformity should be perceived less as a smoothing effect, and more as a pressure to avoid shifting rapidly between content of different information densities, as exemplified by the two sentences earlier in this section. The authors investigate this, among others, by exploring the effect of varying the context window size on the change of variability of the computed information rate, and conclude strongly that the global (smoothing) interpretation wins over the local one. This suggests that the initial explanation of UID in this post as maximizing the use of a metaphorical noisy channel, is, indeed, the most fitting one. As such, language users have a preference for smoothing over the entire language, rather than a on a sentence or phrase-level. A Surprisal-Duration Trade-Off Across and Within the World’s Languages So far we’ve delved into text processing and explored some interesting findings about reading time and linguistic acceptability data. We have also established that linguistic users have a preference for smoothing information across the entire language, as opposed to on a phrase, sentence, or document-level, confirming the most popular interpretation of UID as maximizing the use of a metaphorical noisy communication channel. The second paper summarized in this post investigates UID from a slightly different angle, speech production, and addresses an important question about the linguistic universality of UID interpretations by extending the findings to a corpus of 600 different languages. Basically, this paper explores the relationship between informational content and speech production rate, and defines information content, i.e. surprisal, as the negative log probability of one individual sound, a so-called phone, given a sequence of speech sounds. The hypothesis is that, assuming that the channel capacity of human information processing is roughly the same across languages, it can be expected that languages place various cognitive constraints on how humans smooth the signal over the phonetic inventory of a language in speech production. Namely, the authors demonstrate strong evidence for a trade-off between the surprisal of a phone and the time it takes to produce it; in other words, speakers slow down significantly when producing highly surprising speech sounds, and speed up on more predictable ones. They demonstrate that this condition holds within 319 languages, in which more surprising phones are pronounced with longer duration. Furthermore, they demonstrate that word-initial phones, i.e. sounds occurring at the beginning of a word, are on average more surprising. For data, the authors use a phone-aligned corpus containing readings of the Bible in 600 languages spread over 70 language families. The data contains 20 hours of spoken word on average per language, making their findings the “most representative evidence for UID to date”. The surprisal estimates were computed by a phone-level LSTM model, and the phone durations were given by automatically generated alignments in the running text. The most interesting finding of this paper is that this condition holds not only within one given language, but across languages. Understanding what this means requires a bit of understanding about language universals. Within the world’s languages, there are some tendencies with respect to the features shared by many, although not all, languages; for instance, with respect to phonetics, most languages contain syllables constructed of vowels and consonants where the vowel is the nucleus of the syllable, and most languages contain nasal consonants, [m], [n], or [ng]. Thus, their finding implies that languages containing more universally surprising phones compensate by lengthening the duration of their utterances. Furthermore, the authors didn’t find a single instance in which the oppoosite effect was true; i.e. where shorter phone durations were linked to higher surprisal, and they demonstrate through controlling for several potential confounding variables that this analysis holds remarkably well across languages. Concluding remarks Each of the summarized papers provide evidence of varying strength for the existance of UID and, as such, make very important contributions to computational psycholinguistics. An area of research that might be interesting to expand upon is how well the findings for language comprehention and linguistic acceptability holds across non-Indo-European languages, in order to avoid Anglocentric biases. However, this might be challenging because these languages are generally under-represented when it comes to resources and language models. I hope that through this post, I have given you a better understanding of computational psycholinguistics and how natural language processing is used to further research into the nature of linguistic data. I hope that I have concvinced you that this is not only interesting to a particular subset of academics dealing with linguistics and cognitive science, which one might assume given that this topic does not have immediate industrial applications, but also has interest more generally in the NLP space. Finally, I hope that the journey along this admittedly somewhat dense river of information has been mostly gentle, smooth sailing, and that you at least enjoyed the rock maneuvering along the way.]]></summary></entry></feed>