CTCORPUS

User Guide

Everything you need to work with NKKM. Explore the tools and methodologies behind the Crimean Tatar National Corpus.

Introduction

The National Corpus of the Crimean Tatar Language (NKKM) is a research database of digitized texts. It is used for studying the language, writing dictionaries, and developing language technologies. This guide explains the core functionalities available in our web interface.

NKKM Interface Mockup

N-grams

N-grams are sequences of N words or characters. In NKKM, you can extract the most frequent word combinations (bigrams, trigrams, etc.) to analyze common collocations and syntactic patterns.

Bigram Example menim adım → 1,240 hits
Trigram Example Qırım Tatar tili → 890 hits

Concordance

The Concordance tool allows you to see every occurrence of a word or phrase in its original context (KWIC - Key Word In Context).

Concordance Tools Table

Feature Description
info Result details View metadata like author, year, and source for each result.
filter_list Filter Refine results by text type, time period, or morphological tags.
visibility View options Adjust context length and highlight colors.
download Download Export search results in CSV, Excel, or PDF formats.
settings Change criteria Update your search query without leaving the results page.
casino Get a random sample Select a subset of results for unbiased qualitative analysis.
shuffle Shuffle lines Randomize the order of concordance lines.
sort Sort Sort results by left context, right context, or metadata.
bar_chart Frequency Calculate frequencies of the searched term across different categories.
link Collocations Identify words that statistically co-occur with your search term.
analytics Distribution Visualize how the term is distributed across the entire corpus timeline.
view_agenda KWIC/sentence view Toggle between keyword-centered view and full sentence view.

GDEX

GDEX (Good Dictionary EXamples) is an algorithm that automatically identifies the best representative sentences for a word. It prioritizes sentences that are clear, not too long, and grammatically complete—perfect for lexicographers and language learners.

Wordlist

Generate lists of words based on frequency or alphabetical order. You can filter wordlists by part of speech, length, or specific character patterns to find unique linguistic data points.

Keywords

Keyword analysis helps identify words that are unusually frequent in a target text compared to a reference corpus. This technique highlights the distinctive themes and terminology of specific sub-corpora.

Text types analysis

This tool provides statistical breakdowns of search results based on text metadata. Analyze word usage across genres (e.g., fiction vs. news), dialects, or chronological periods to observe linguistic evolution and stylistic variation.

Video Tutorials

Below you will find video tutorials demonstrating the key features of the corpus on the Sketch Engine platform.