Skip to content

FIX: Preserve text columns when loading human-labeled CSVs - #2980

Open
Garyouki wants to merge 2 commits into
microsoft:mainfrom
Garyouki:fix/human-labeled-csv-text
Open

Garyouki wants to merge 2 commits into
microsoft:mainfrom
Garyouki:fix/human-labeled-csv-text

Conversation

@Garyouki

@Garyouki Garyouki commented Oct 4, 2026

Copy link
Copy Markdown
Contributor

Description

Closes #2978.

HumanLabeledDataset.from_csv() loaded response 00123 as 123 and objective 00456 as 456 because pandas inferred numbers before the loader converted them back to strings. Read response, objective, harm-category, and data-type columns as strings in both the UTF-8 and Latin-1 paths. Human-label parsing and missing-value validation retain their existing behavior.

Tests and Documentation

  • Added 16 CSV-loading cases spanning numeric/boolean-looking responses, objective/harm datasets, and UTF-8/Latin-1. Assertions cover original and converted response text, objective/category text, and typed human labels.
  • Before the fix: 16 failed, 73 passed. After: 89 passed (pytest tests/unit/score/test_human_labeled_dataset.py -q --no-cov).
  • Ruff lint/format, targeted ty check, async-suffix and reST-role checks, and git diff --check passed.
  • No public API change. Full repository tests were not run.

@hannahwestra25 hannahwestra25 self-assigned this Oct 8, 2026

@hannahwestra25 hannahwestra25 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

thanks for contributing !

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

HumanLabeledDataset.from_csv silently coerces numeric response and objective text

2 participants