Machine-learning project that estimates football player market values (Transfermarkt-style data) and highlights under- and over-valued players. Includes a Streamlit dashboard for exploration, descriptive analysis, talent search, and live predictions with position-specific XGBoost models.
Deskriptive Analyse — distributions of foot preference, position, age, and market value:
Talentsuche — filters and players where predicted value exceeds market value:
Marc Sieber, Stella Sun, Linda Fuchs, Eliane Elsässer, Jonas Vogel
- Scrape player pages from Transfermarkt (Selenium workers, configurable via
scraping/config.yaml) - Clean & merge scraped CSVs into modeling tables
- Train & evaluate “simple” and “extensive” models (notebooks under
analysis/) - Serve insights via Streamlit (
streamlit/app_v0.py)
Public demo (may be stale): Streamlit Cloud app
| Path | Purpose |
|---|---|
scraping/ |
manager.py, worker.py, raw scraped_data/, cleansing notebooks |
analysis/ |
Jupyter notebooks — EDA, preprocessing, training, performance CSVs |
data/ |
Processed CSVs and prediction outputs for the app |
models/ |
Pickled XGBoost models (simple-model-xgb.pkl, goalkeeper variant, extensive models) |
streamlit/ |
Dashboard entrypoint and slim requirements.txt |
Paths in app_v0.py are relative to the repository root (data/, models/).
cd FootballPredictions
python3 -m venv .venv
source .venv/bin/activate
pip install -r streamlit/requirements.txt
streamlit run streamlit/app_v0.pyOpen http://localhost:8501.
Required artifacts (committed in this repo):
data/df_model_full_merge.csv,df_simple_model_*_streamlit.csvmodels/*.pkl
Scraping is ** brittle** (site layout, chromedriver, legal/ToS). For portfolio purposes, rely on committed data/ and models/.
High-level steps documented in notebooks:
scraping/manager.py→ raw files inscraping/scraped_data/scraping/data_cleansing.ipynb→ cleansed tablesanalysis/empty_values.ipynb→data/df_clean.csv- Preprocessing notebooks per model →
data/df_*_model_*.csv - Model notebooks →
data/*_results.csvandmodels/*.pkl analysis/model_predictions_full_merge.ipynb→df_model_full_merge.csv
Transfermarkt data was used for academic analysis. Do not use scrapers against third-party sites without respecting terms of service and robots rules. This repository is for demonstrating methodology, not for production scraping.
MIT License — University of St.Gallen course project; Transfermarkt data subject to their terms of use.

