trstop

Turkish Stop Words Türkçe Dolgu Sözcükleri In this repository I put Turkish stop words that is contained in the first 10 thousand words with the highest frequency. In order to test the new candidate words in future, I add a small python script, and a 10 thousand item word list with highest frequency. At https://github.com/sgsinclair/trombone/blob/master/src/main/resources/org/voyanttools/trombone/keywords/stop.tr.turkish-lucene.txt are some Turkish stop words. However, some stop words in that list do not belong to the ten thousand highest frequency words.

In order to use the module:

import trstop

print(trstop.is_stop_word(parameter))

Contributors:

Ahmet Aksoy
Toprak Öztürk

Bu depoya en sık kullanılan 10 bin Türkçe sözcük listesinde yer alan dolgu sözcüklerini ekledim. Dolgu sözcükleri (stop words), sık kullanılan, ama iptal edildiklerinde ayrıldıkları cümlenin anlamında önemli değişiklikler oluşturmayan sözcüklerdir.

"Stop words" terimine karşılık "dolgu sözcükleri" terimini kullandım. Daha iyi bir seçenek varsa, değiştirmeye hazırım. Depoya eklediğim "turkce-stop-words-dict.py" betiğini, ileride listeye yeni sözcükler eklemek istediğimizde kullanım sıklığını denetlemek amacıyla kullanabiliriz.

https://github.com/sgsinclair/trombone/blob/master/src/main/resources/org/voyanttools/trombone/keywords/stop.tr.turkish-lucene.txt adresinde de bazı dolgu sözcükleri listelenmiş. Ancak buradaki bazı sözcükler ilk on bine girecek kadar yoğun frekansa sahip değil.

Modülü kullanmak için:

import trstop

print(trstop.is_stop_word(parametre))

Projeye katkıda bulunanlar:

Ahmet Aksoy
Toprak Öztürk

Son güncelleme: 29.06.2018

Turkish Stop Words Türkçe Dolgu Sözcükleri

Related tags

Overview

trstop

In order to use the module:

Contributors:

Modülü kullanmak için:

Projeye katkıda bulunanlar:

Owner

Ahmet Aksoy

CrossNER: Evaluating Cross-Domain Named Entity Recognition (AAAI-2021)

Prompt-learning is the latest paradigm to adapt pre-trained language models (PLMs) to downstream NLP tasks

A Persian Image Captioning model based on Vision Encoder Decoder Models of the transformers🤗.

Weird Sort-and-Compress Thing

This repository contains Python scripts for extracting linguistic features from Filipino texts.

Sequence model architectures from scratch in PyTorch

Simple text to phones converter for multiple languages

Reproduction process of BERT on SST2 dataset

A telegram bot to translate 100+ Languages

Finds snippets in iambic pentameter in English-language text and tries to combine them to a rhyming sonnet.

HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis

A Telegram bot to add notes to Flomo.

Dope Wars game engine on StarkNet L2 roll-up

Repository for the paper: VoiceMe: Personalized voice generation in TTS

AMUSE - financial summarization

The ibet-Prime security token management system for ibet network.

A Practitioner's Guide to Natural Language Processing

A python wrapper around the ZPar parser for English.

Kashgari is a production-level NLP Transfer learning framework built on top of tf.keras for text-labeling and text-classification, includes Word2Vec, BERT, and GPT2 Language Embedding.

Every Google, Azure & IBM text to speech voice for free