Repository navigation
fix(human_labeled_dataset): preserve # inside unquoted CSV cells - #2983
Chen Yufeiyang (feiiiiii5) wants to merge 2 commits into
Conversation
| with open(csv_path, encoding="utf-8", errors="ignore") as f: | ||
| for line in f: | ||
| if line.lstrip().startswith("#"): | ||
| skiprows += 1 | ||
| else: | ||
| break |
There was a problem hiding this comment.
| with open(csv_path, encoding="utf-8", errors="ignore") as f: | |
| for line in f: | |
| if line.lstrip().startswith("#"): | |
| skiprows += 1 | |
| else: | |
| break | |
| with open(csv_path, encoding="utf-8-sig", errors="ignore") as f: | |
| for line in f: | |
| stripped = line.strip() | |
| if stripped.startswith("#") or not stripped: | |
| skiprows += 1 | |
| else: | |
| break |
lstrip() doesn't strip U+FEFF, so a UTF-8-BOM file (Excel's default) leaves skiprows=0 and the metadata line becomes the header — ValueError: Required column 'assistant_response' is missing. Worked before, since comment="#" ran after BOM handling. A leading blank line fails the same way. plus adding unit test
There was a problem hiding this comment.
| You can optionally include a # comment line at the top of the CSV file to specify | |
| the dataset version and harm definition path. Only leading # lines are treated as | |
| comments; a # anywhere else is preserved as literal data. The format is: |
The scan for leading "#" comment lines read the file as utf-8, so a byte-order mark left the first line looking like data and a blank line stopped the scan, which dropped the metadata comment into the header row. Both reads now use utf-8-sig, and the version parse reads the first line the same way, so a BOM file's comment line is recognized as the comment line it is.
|
Good catch on both, and reproduced before changing anything — you were right about the cause. The BOM/blank-line bug. Confirmed on the branch head: a file written as Fixed in skiprows = 0
with open(csv_path, encoding="utf-8-sig", errors="ignore") as f:
for line in f:
stripped = line.strip()
if stripped.startswith("#") or not stripped:
skiprows += 1
else:
breakTwo unit tests: The docstring. Added your wording: "Only leading |
Fixes #2974
HumanLabeledDataset.from_csvpassescomment="#"to pandas, which treats every#anywhere in the row as the start of a comment — so a hash inside an unquoted cell ("mentions C#",# Heading, a leading#in an "assistant_response") silently drops fields or the entire row.csv.writer/DataFrame.to_csvdo not quote fields just because they contain#.Skip only the leading
#comment lines (already parsed for metadata) withskiprowsand remove the pandascommentargument, so the rest of each row is parsed verbatim.Test:
Fails on
main(1 entry, objectivementions Cin the issue's repro), passes on this branch (74 passed).ruff check/ruff format --checkclean on the touched files.