Tabiya Matching Engine is a matching service that recommends occupations and job opportunities for users based on skills, preferences, and market signals.
The repository contains:
backend: FastAPI service for scoring and recommendation APIs.frontend: React application for interacting with matching outputs.- shared resources and scripts for benchmarking, diagnostics, and operational maintenance.
The backend supports multi-user requests, Mongo-backed job retrieval, and configurable scoring behavior for both quality and latency tuning.
- User-to-opportunity matching with ranked recommendations.
- User-to-occupation matching for broader career pathways.
- Skill gap recommendations to improve future match potential.
- Configurable scoring and response thresholds via environment variables.
Default scoring mode is multiplicative (SCORING_MODE=multiplicative):
S_total = U_hat × P_hat
Where:
U_hatcaptures utility from skills and preferences.P_hatcaptures success propensity (gate, essential fit, readiness, market opportunity).
Legacy additive mode is also available (SCORING_MODE=additive) for controlled comparisons.
Primary endpoint:
POST /match— accepts one or more users and returns:opportunity_recommendationsoccupation_recommendationsskill_gap_recommendations
Hybrid diagnostic / alternate ranking:
POST /match_v2— sameMatchRequestbody shape asPOST /match(JSON array); loads all active jobs from Mongo without the per-user location prefilter used byPOST /match(JOBS_RETRIEVAL_FILTERis effectively bypassed here so hybrid indexes match unrestricted batch runs, e.g. CLI--mongo-all-active). Returnshybrid_recommendationsranked by BM25 × embedding‑cosine pool fused scores (optional query:fusion_top_k,alpha_on_cosine). Does not compute occupations or the full SkillScorer /p_hatstack.x-api-keyis not required on this route for now (unlike/match).
The language a deployment matches in is configured with TARGET_LANGUAGE (see
Languages), not per request. Skill matching itself is language-neutral, so a
Spanish posting matches a Spanish profile either way.
Interactive API docs are available at http://127.0.0.1:8000/docs when the backend is running.
cd backend
python -m venv venv
source venv/bin/activate
pip install -r requirements.txt
./setup.sh
uvicorn app.main:app --reloadcd frontend
npm install
npm run devEach deployment is configured for one language with TARGET_LANGUAGE (en | es, or a
locale spelling like AR-es / es_AR / spanish). Requests carry no language. The
important thing to understand is which half of the pipeline is language-neutral and which
is not.
Skill matching is language-neutral. Both sides resolve skills by label into the internal id space of the embedding artefact. Every enabled language's taxonomy label pack is loaded into that one resolver and mapped onto the same canonical ids, so a Spanish job posting matched against a Spanish user profile scores through exactly the same vectors as the English equivalent — with nothing on the request, and with no Spanish retrain.
That works because skill IDs are per-taxonomy-locale but UUIDHISTORY's oldest entry is
not: it is identical across locales for all 13,896 skills. The packs are joined on it at
load time (app/services/skill_label_packs.py).
Text scoring and display are not. These follow TARGET_LANGUAGE:
| What | Where |
|---|---|
Cross-encoder checkpoint (stage-2 rerank on /match_v3, /match_v4) |
cross_encoder_model per language; CROSS_ENCODER_MODEL_NAME_<LANG> overrides |
| BM25 / hybrid stopwords | stopwords per language |
| Labels echoed back in the response | SkillScorer.display_labels(language) |
| Occupation database labels | resources/occupations/<lang>/, falling back to en |
An unset TARGET_LANGUAGE means en; an unregistered value falls back to en with a
warning at startup rather than failing the deployment.
# An Argentina deployment: Spanish postings + Spanish profiles, Spanish-capable reranker
TARGET_LANGUAGE=es uvicorn app.main:appOn Cloud Run it is one variable per stack: TARGET_LANGUAGE in the stack's GitHub
environment (vars.TARGET_LANGUAGE), passed through iac/backend/env_vars.py. Leave the
SKILLS_CSV_PATH / SKILL_GROUPS_CSV_PATH / SKILL_HIERARCHY_CSV_PATH /
OCCUPATION_JSON_PATH vars empty — each one pins every language to a single file (see
iac/backend/.env.example).
Registered languages live in backend/app/languages/ (en_config.py, es_config.py);
LANGUAGE_REGISTRY in __init__.py is the only list to edit.
- Add the code to
LANGUAGE_REGISTRYinbackend/app/languages/__init__.py. - Copy
es_config.pyto<code>_config.py; set its locales, cross-encoder checkpoint and stopwords. - Build its taxonomy label pack from a taxonomy CSV export:
cd backend
python -m scripts.build_language_taxonomy --taxonomy-dir <export-dir> --language <code>The script validates the columns the resolver reads by name and — the part that matters — reports how much of the pack joins onto the canonical id space. Anything that does not join has no embedding row, so labels resolving to it would be silently dropped at match time; that almost always means the two packs came from different taxonomy releases.
tests/unit/test_language_support.py guards the invariant: every pack must join onto the
canonical id space, and a Spanish label must resolve to the same id as its English
counterpart.
ENABLED_LANGUAGES limits which packs are loaded (default: all — it is a CSV parse, not a
model load). The canonical language is always included; it defines the id space.
Backend runtime settings are managed through backend/.env (see backend/.env.example).
Key settings include:
- data source and retrieval controls (Mongo collection, retrieval filters, projection, warmup)
- language defaults (
TARGET_LANGUAGE,ENABLED_LANGUAGES,CROSS_ENCODER_MODEL_NAME_<LANG>) - scoring mode and weights
- top-k response sizes
- response skill thresholding (
MATCH_RESPONSE_SKILL_MIN_SCORE)
If MATCH_RESPONSE_SKILL_MIN_SCORE is not set, it falls back to GATE_SIMILARITY_THRESHOLD.
Every matching request — /match, /experiments/v2/match, /experiments/v3/match, /match_v4,
/experiments/v5/match (and any future /match_*) — can be traced to Langfuse,
with the same layer the llm-reranker and compass-connect use (backend/app/observability/). It is
off by default; a deployment with no Langfuse keys behaves exactly as before.
What one trace holds:
| Observation | What it shows |
|---|---|
| root, named after the route | request id, pseudonymous user_id (single-user requests; batches list user_ids), query params, HTTP status, embedding totals |
retrieval |
Mongo find / build and occupation-cache timings, job and occupation counts |
embedding → embed_content |
one Langfuse embedding per Gemini call: model, dimensionality, token usage, attempts, retries, per-attempt latency, failure |
shortlist, rerank |
stage-1 cosine (or BM25 × cosine for v2) and the cross-encoder, per corpus (jobs / occupations) |
preference_scoring, formatting, skill_gaps |
u_hat × p_hat scoring, row building, skill-gap analysis |
Stage names are the same on every route, so a Langfuse dashboard of observation latency
(p50 / p95 / p99) grouped by name gives per-stage percentiles; filter by the route:<path> tag for a
single route and by environment for a deployment (IaC sets it to the Pulumi stack).
- Find a request: traced responses carry
X-Request-ID(a client-sent one is echoed) andX-Trace-ID— paste the trace id into Langfuse. Or search by the user's id. - Failures and retries: traces are tagged
error(5xx / exception),client_error(4xx),embedding_retriedandembedding_failed. - Embedding spend: Langfuse prices each
embed_contentfrom its token usage and the model's price. Gemini'sembed_contentreports no tokens, so a traced call runscount_tokensin parallel with it and waits at mostTOKEN_COUNT_WAIT_S(0.2 s) after the embedding returns; if the count is not in by then, that call records no usage rather than a guess. If the Langfuse project has no price forgemini-embedding-001, add it once under Settings → Models (matchgemini-embedding-001, input price per token); cost then shows per trace and aggregates per environment. - Privacy: no request or response body is recorded — only counts, timings, error class names,
the request id and the pseudonymous user id. The text sent to Gemini (the jobseeker's skills) is
exported only with
"recordEmbeddingInput": true, masked first. Vectors are never exported.
MATCHING_ENABLE_TRACING=1
MATCHING_LANGFUSE_HOST=https://cloud.langfuse.com
MATCHING_LANGFUSE_PUBLIC_KEY=pk-lf-...
MATCHING_LANGFUSE_SECRET_KEY=sk-lf-...
MATCHING_TRACING_ENVIRONMENT=dev # IaC: defaults to the stack name
MATCHING_TRACING_CONFIG={"sampleRate": 1.0, "recordEmbeddingInput": false}In GitHub, set MATCHING_ENABLE_TRACING (and optionally MATCHING_LANGFUSE_HOST,
MATCHING_TRACING_CONFIG) as environment variables and MATCHING_LANGFUSE_PUBLIC_KEY /
MATCHING_LANGFUSE_SECRET_KEY as environment secrets. The stderr timing blocks
(app/match_timing_log.py) are unchanged.
Cloud Run deployment is supported through:
backend/build-and-deploy.sh
Example:
cd backend
./build-and-deploy.sh <project-id> <env-vars-yaml>