Professional Services & Staffing · Taxonomy Unification AI
Taxonomy Unification AI
Using ML and NLP to make recruiting data more consistent across systems
The Client · A NASDAQ-listed staffing and professional-services firm

Overview
A NASDAQ-listed staffing and professional-services firm engaged Taller to build an AI taxonomy engine for normalizing skills and job descriptions across its recruiting data.
The Problem
The client’s recruiting data was spread across many sources: job descriptions, candidate profiles, resumes, and internal recruiting workflows. But the language describing it was inconsistent. The same skill could be described several different ways, and the same role title could mean different things depending on the employer, industry, seniority, or business context. The titles "developer," "software engineer," "full-stack engineer," and "application engineer" might overlap heavily in one setting yet describe very different profiles in another.
That inconsistency created a structural problem for AI matching, as even strong matching models produced weak results when the underlying language was fragmented. If skills were not normalized, if role families were not represented consistently, and if job descriptions were not converted into a common taxonomy, the system could not reliably compare one job to another, or one candidate to one job. The client needed an AI-driven taxonomy layer that could understand recruiting language, group related skills and jobs, normalize messy descriptions, and produce a consistent picture of demand across the business.
The Solution
Taller built the engine in four layers, combining natural-language processing (NLP), unsupervised learning, supervised classification, deep learning, and custom data-engineering pipelines.
The first layer handled skill detection. Taller built NLP pipelines to extract skills from job descriptions and candidate profiles, catching both the explicitly named skills and related terms expressed in different wording, so the system could tell that two differently worded descriptions were asking for the same underlying capability.
The second layer handled grouping and normalization. Taller used clustering and nearest-neighbor similarity analysis (unsupervised learning techniques that group items by similarity without predefined categories) to surface natural clusters in the data: skills that commonly appeared together, descriptions belonging to the same functional area, and role families similar in meaning even when their titles were not.
The third layer built the taxonomy itself. Taller used the models' output to construct a normalized skills-and-jobs taxonomy tailored to the client’s recruiting business, learning from the client’s own market data rather than relying on generic public taxonomies. Job descriptions were then transformed into structured representations using this taxonomy, so they could be compared, searched, classified, and matched far more accurately.
The fourth layer connected the taxonomy to recruiting intelligence. Once jobs and skills were normalized, the system could classify roles by business area, detect demand patterns, improve job-order matching, and support candidate scoring. The taxonomy became a service other recruiting workflows could draw on.
This work predated the current LLM wave, built entirely through custom NLP and ML research rather than general-purpose language models.
The Impact
The taxonomy engine delivered:
More importantly, it gave the client a reusable AI foundation for matching, classification, job analysis, and recruiting intelligence. Jobs and skills that had been fragmented across inconsistent language could now be represented in one normalized structure that downstream systems could use.
reduction in data discrepancies
improvement in candidate-matching accuracy
increase in recruiting efficiency


