(This blog is based on a joint research and publication in collaboration with David Alfter, Therese Lindström Tiedemann, Maisa Lauriala and Daniela Piipponen)
Aktuellt från Språkbanken
Så kan språkteknologi stärka små språk
Grierson’s “Linguistic Survey of India” as open-access digital data resource for studying languages of South Asia
Lars Borin, Anju Saxena, Shafqat Mumtaz Virk, Bernard Comrie
Projekt om svenska språkmodeller får forskningmedel
Snart premiär för ny svensk korpus
Språkbanken bidrar till nya sätt att tillgängliggöra KB:s samlingar
Spana i dialektkartan!
En syntaktisk beskrivningsmodell för modern svensk text
Sverige har en relativt lång tradition av att skapa en typ av korpus som brukar kallas trädbank. En trädbank är en samling texter som har annoterats (märkts upp) med ordklasser och syntaktisk struktur. Den syntaktiska strukturen för en mening kan ritas upp så att den liknar ett träd.
Korp searches in Second Language data
Korp offers a lot of different corpus collections for various types of search (and research). Swedish as a Second Language (L2) is one of the subcategories of the language that can be studied with the help of Korp. At the moment, Korp provides access to five L2 corpora through its interface:
The five lives of Talbanken
This post is about Talbanken, one of the most widely used and important Swedish corpora. There exist at least five versions of this treebank, and the purpose of this post is to reduce ambiguity of the name "Talbanken", which sometimes leads to confusion.
Om ordklasser för svenska språket
Ordklassindelning används i många språkteknologiska verktyg därför att det är ett sätt att skilja mellan olika användningar av ett ord. Genom ordklasserna kan man enklare söka efter liknande ord och uttryck i stora textmängder, eller skapa en ny text med liknande form.
En topic modell bland andra – En data-intensiv forskningsmetodologi 3
I det första avsnittet av denna bloggserie pratade vi om den data-intensiva forskningsmetodologin och gav en övergripande bild.
Text som forskningsdata – En data-intensiv forskningsmetodologi 2
Detta blogginlägg är en uppföljning av ett tidigare som inleds med en beskrivning av en data-intensiv forskningsmetodologi, börja gärna med den.
Common Pitfalls in the Development of ICALL Applications
This blog is a piece of opinion where I sketch the process of developing NLP-based applications for second language learning and look at the process from the point of view of typical (mis)conceptions and challenges, as I have experienced them. Are we over-trusting the potential of NLP?
En data-intensiv forskningsmetodologi 1
I en värld där AI tar en allt större plats har datadriven forskning blivit orden på allas läppar. I det här blogginlägget tänkte jag prata lite om vad det innebär att forska med hjälp av stora mängder textdata, primärt inom humaniora.
A multilingual annotated corpus of world's natural language descriptions
Shafqat Mumtaz Virk, Harald Hammarström, Markus Forsberg, Søren Wichmann
Zipfs lag på svenska
Zipfs lag, uppkallad efter den amerikanske lingvisten George Kingsley Zipf, säger att ett ords frekvens är omvänt proportionellt mot dess plats i en frekvenslista. Vad innebär det?
Meaning through sensory data
Recently, we have seen a surge of methods that claim to embed meaning from textual corpora. But is that possible? Can text really reveal meaning, and if so, can current NLP methods detect it? Can our methods, as they some times claim, understand?
The Gothenburg H70 birth cohort studies and the digital assessment of neuropsychological tests
A comment often received by the reviewers of manuscripts to scientific conferences and journals is one about the representative sample under scrutiny and whether there are any solid arguments for accepting that the population characteristics, and particularly the features extracted from the empir
Argumentation Mining
What if you could find all arguments in a text without having to read it? Or, what if you could search a database for a controversial topic and immediately get arguments for and against it, gathered from text all around the internet?