PROOF FIRST — QA EVIDENCE PACK TEST PLAN Classification Independent portfolio exercise. This is not client work, a security certification, or a new vulnerability claim. System under test - Vyper compiler v0.4.3, commit bff19ea204059290da652854cd634abef10f6c43 - EVM target: Prague for the differential campaigns - Local execution backends: REVM and Py-EVM - Production optimizer modes: none, gas, codesize Objective Detect observable semantic divergence across compiler configurations and local EVM implementations. Compare return values, reverts, state changes, and event logs against independent semantic oracles. Entry criteria 1. The harness unit tests pass. 2. Campaign scripts pass Python syntax checks. 3. The historical public positive control reproduces on Vyper v0.4.1. 4. The corrected positive control passes on Vyper v0.4.3. Test strategy 1. Freeze the observable behavior and comparison schema before generation. 2. Calibrate the detector with the public zero-length concat side-effect defect. 3. Generate bounded cases for each campaign family. 4. Execute every eligible case across two backends and three optimizer modes. 5. Compare each observation with its semantic oracle. 6. Triage mismatches as harness defects, backend differences, known issues, or compiler candidates before making any claim. Campaign results - Base @raw_return cases: 264 clean semantic executions - @raw_return provenance and control flow: 10,230 - Dynamic ABI encoding: 5,424 - Nested dynamic ABI encoding: 4,608 - bytesM bitwise and padding: 12,960 - Stateful expression trees: 1,440 - Typed event side effects and encoding: 864 - Effectful composite literals in events: 1,440 - Total: 37,230 / 37,230 matched their independent semantic oracles Integrity checks - Differential harness tests: 9 / 9 passing - Positive control: vulnerable v0.4.1 behavior detected - Corrected control: v0.4.3 behavior confirmed Exit criteria - Each included observation is captured across all six production configurations. - Any mismatch survives deterministic reproduction and duplicate/known-issue review. - The report states exact counts, exclusions, and negative evidence. Exclusions and limits - Experimental code generation - Public networks and mainnet execution - Protocol-wide security review - Claims that a clean campaign proves the absence of compiler defects Disposition No eligible new vulnerability was found. The result is bounded negative evidence: these specific generated cases produced no candidate mismatch.