fix(ai-file): stop recompressing PPTX window package members
_create_slide_window builds a throwaway package (opened, then unlinked a few
hundred ms later) that references only the requested slides. It replaces two
XML parts and copies the rest. The copy was not a copy: copy(info) carries
the source compress_type over, so source.open() (inflate) piped into
target.open() (deflate) recompressed every member. PowerPoint stores all
members DEFLATED, so there was no exception - the whole package was inflated
and deflated again on every read, producing bytes nearly identical to the
originals. shutil.copyfileobj hid it: nothing in that loop mentions
compression.
Media dominate the bytes and are already compressed - recompressing them shrank the synthetic 16.9MB deck by 1.3% for 0.358s of CPU. Decide per member from the source zip header (no extra I/O): incompressible members are stored, compressible ones (XML) keep deflate at level 1. Blanket ZIP_STORED was rejected because it inflates XML-heavy decks 3x (1.01MB -> 3.16MB), which would just move the cost to scratch writes.
Measured window creation (macOS, 5-slide window): 16.88MB / 336 members 0.410s -> 0.062s (6.6x), window 16.88 -> 16.93MB 30.88MB / 3036 members 0.813s -> 0.220s (3.7x), window 30.88 -> 31.16MB 0.98MB / 2036 members 0.126s -> 0.079s (1.6x), window 0.98 -> 1.03MB
Context: pptx /read measured 3.48s per request under load against PDF's
0.569s, and ~89% of that was this function. Slide-count safety is unchanged -
the window still exists, so python-pptx keeps loading only the requested
slides (peak memory grows with slide count, not file size, and the 20MB file
guard does not bound it).
Tests: 9 new. The regression guard builds a deck with noise-JPEG media and asserts those members are stored in the window; it fails with the fix removed (verified). Full suite 488 passed.
Co-Authored-By: Claude Opus 5 (1M context) noreply@anthropic.com