Skip to content

feat(xlsx): path-based streaming parse

김대순 requested to merge feature/agent-knowledge-pdf-page-parallel into main

Move the xlsx branch above the file_path.read_bytes() shim in read() and extract_metadata(), mirroring the Task 8 PDF migration. openpyxl load_workbook now takes file_path directly (read_only=True streams rows) instead of BytesIO(file_bytes), so XLSX parsing/metadata no longer materializes the full workbook bytes in memory. Adds negative-control tests asserting Path.read_bytes() is never called for XLSX read/metadata.

Merge request reports