Skip to main content
Have a personal or library account? Click to login
Automated Extraction and Classification of IT Job Market Data Using Locally Deployed Large Language Models Cover

Automated Extraction and Classification of IT Job Market Data Using Locally Deployed Large Language Models

Open Access
|Jul 2026

Abstract

Understanding labor market dynamics requires a systematic analysis of job posting data, but manual processing of large volumes of unstructured text remains impractical. This paper presents an automated workflow for collecting, filtering, and semantic enrichment of IT labor market data from Romanian recruitment platforms. We collected 56.062 publicly available job postings from BestJobs.ro over 244 days, using automated web data extraction techniques. The raw data went through a multi-stage processing workflow, using locally implemented large language models (Llama3.3:70b, Qwen2.5:14b and Qwen2.5:7b-instruct) through the Ollama platform, for category filtering, occupational classification according to the ESCO taxonomy, skills extraction, and semantic skills mapping. The flow reduced the initial data set to 2,152 validated job postings, classified into 75 ESCO occupations, with 2,797 unique skills extracted and mapped to 881 ESCO skill concepts. The methodology demonstrates that locally deployed open-weight LLMs can effectively replace cloud-based APIs for large-scale information mining tasks, while maintaining data confidentiality and eliminating per-query costs. Our approach provides a reproducible framework for processing labor market data, which can be adapted to other domains and languages.

Language: English
Page range: 4027 - 4037
Published on: Jul 22, 2026
In partnership with: Paradigm Publishing Services
Publication frequency: 1 issue per year

© 2026 Lucian VILCEA, Liviu-Adrian COTFAS, Ioana IOANĂȘ, George-Cristian TĂTARU, published by Bucharest University of Economic Studies
This work is licensed under the Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 License.