Location : Faculty of Mining and Geology, University of Belgrade

TESLA: From Text to Knowledge Graphs

Named Entity Recognition and Linking to Wikidata, Relation Extraction, and Knowledge Graph Construction for Serbian

Workshop Description

Within the TESLA project, advanced resources and tools have been developed for Named Entity Recognition (NER), Named Entity Linking (NEL), and Semantic Relation Extraction (RE). Although these tasks are often studied independently, their integration enables the automatic construction of knowledge graphs from unstructured texts.

This workshop presents a complete workflow for transforming text into structured knowledge using language resources for Serbian. Participants will be introduced to the TeslaNER+ corpus, which contains more than 150,000 sentences annotated with named entities linked to Wikidata, as well as methods for relation extraction based on local grammars, finite-state transducers (FSTs), and large language models (LLMs).

Special attention will be devoted to linking entities to Wikidata, working with QID identifiers, retrieving additional information through SPARQL queries, and enriching entity descriptions using linked open data resources.

In the final part of the workshop, participants will learn how to automatically construct knowledge graphs from extracted entities and relations, represent them as RDF triples, and visualize them using graph exploration and analysis tools.

Learning Outcomes

By the end of the workshop, participants will be able to:

  • recognize named entities in text;
  • link entities to Wikidata (Named Entity Linking);
  • use Wikidata QIDs and SPARQL queries;
  • extract semantic relations between entities;
  • work with relation extraction models;
  • generate RDF triples in the form subject–predicate–object;
  • construct and enrich knowledge graphs;
  • visualize and analyze knowledge graphs;
  • apply knowledge graphs in information retrieval, question answering, and digital humanities.

Hands-on Session

Participants will work with real-world examples from different domains, including literature, biographies, history, and news, using authentic data from Serbian language corpora.

The complete processing pipeline will be demonstrated:
Text → NER → NEL → RE → RDF Triples → Knowledge Graph

Target Audience

The workshop is intended for students, linguists, lexicographers, researchers in natural language processing, digital humanities scholars, librarians, information scientists, and anyone interested in semantic technologies, linked open data, and knowledge graphs.

Prerequisites

Basic familiarity with text data and corpora is recommended. No prior experience with Wikidata, RDF, or knowledge graphs is required.

Organizers

  • Ranka Stanković, University of Belgrade – Faculty of Mining and Geology
  • Milica Ikonić Nešić, University of Belgrade – Faculty of Philology
  • Saša Petalinkar, University of Belgrade – Faculty of Mining and Geology
  • Olivera Kitanović, University of Belgrade – Faculty of Mining and Geology