Research, written in the open.
Long-form engineering research from inside newmind AI. Models, corpora and code are released alongside each piece.
Training GLiNER2.5 for the Turkish KVKK Law PII task
Three open models for the KVKK personal-data task in Turkish, the synthetic corpus they learned from, and the code that produced both. Every value, its category and the person it belongs to, from one pass over a document.
Read the article →