Q-A to Q-H | 2026-10-08 | Controlled English based on ASD-STE100 Issue 9
The guide explains eight static hypotheses. It covers 35 Qwen2.5-7B entries and two Qwen3.5-4B embedding entries.
These IDs are not CAND numbers. No item has measured whole-model performance.
Figure 1. Compiler boundaries. This figure is a schematic, not a timing chart.Figure 2. Evidence and status. This figure is a schematic, not a timing chart.
Blue means observed code. Green means an arithmetic relation. Dashed amber means a hypothesis. Dashed gray means an unverified boundary.
The review uses saved artifacts. No new compiler, GPU or Nsight run supports this guide.
Common SCF, arithmetic or LLVM loop analysis / boundary open
Range of a small tile loop
Q-A | Shifted contiguous gather
7B: 0 | 4B embedding: 0 | Common MLIR Vector and Vector-to-LLVM
Figure 3. Q-A: The gather path. This figure is a schematic, not a timing chart.Figure 4. Q-A: Addresses and the proposed recognition. This figure is a schematic, not a timing chart.
Computation
The embedding kernel copies one table row. Each thread reads eight adjacent half values.
Observed code
The IR keeps a shifted gather. Offline SASS has eight U16 loads and packing instructions.
Source hypothesis
The common matcher identifies zero-based sequences. A source extension can identify the shifted sequence.
Meaning conditions
The transform needs unit stride and valid cast ranges. It must keep the mask and pass-through behavior. Negative-token handling must stay unchanged.
Open boundary
The first gather creation and fold positions remain unknown. The generic draft has no execution result.
Figure 5. Q-B: The row address changes form. This figure is a schematic, not a timing chart.Figure 6. Q-B: The alignment relationship. This figure is a schematic, not a timing chart.
Computation
The output projection reduces a dot product across 28 warps. Each thread reads four adjacent A values.
Observed code
O3 changes the row address to x minus its remainder. Four scalar A loads keep only align 4.
Source hypothesis
KnownBits treats the operands separately. A bounded relation can show the zero low bits of the row address.
Meaning conditions
The analysis must keep freeze, undef and poison semantics. It must also respect the divisor and integer width.
Open boundary
The first missed alignment query has no trace. The A base still has align 64.
35 A base: %1. 61 A base: %0. Scalar loads: %45, %47, %49, %51.
35 A loads: 03c0, 03e0, 0400, 0420. 61 A loads: 03b0, 03d0, 03f0, 0410.
Q-C | Alignment fact and instruction sinking
7B: 15, 37, 38, 743 | LLVM InstCombine
Figure 7. Q-C: The alignment fact disappears. This figure is a schematic, not a timing chart.Figure 8. Q-C: Source behavior and the open trace. This figure is a schematic, not a timing chart.
Computation
These projections reduce dot products. The A input supplies adjacent half values to each thread.
Observed code
The A alignment assume disappears in O3. The moved A address then supplies loads with align 2.
Source hypothesis
Instruction sinking can remove droppable uses. The source behavior agrees with the observed change.
Meaning conditions
A source change must keep dominance and execution-path rules. The historical patch does not establish a Qwen speedup.
Open boundary
The exact Qwen sink step has no trace. This item continues G01. It is not a new independent candidate.
15: Q projection. 37: up projection. 38: gate and SiLU. 743: lm_head.
The surviving %27 assume in dispatch 38 refers to another epilogue input, not A.
Q-D | Split and rebuilt address
7B: 13, 31 | 4B embedding: 1 | LLVM InstCombine or common MLIR arithmetic
Figure 9. Q-D: Split and rebuild the address. This figure is a schematic, not a timing chart.Figure 10. Q-D: The same address with fewer terms. This figure is a schematic, not a timing chart.
Computation
The kernels convert or scale eight values per thread. Their addresses split a flat index into quotient and remainder.
Observed code
Casts and OR separate the related terms. SASS keeps quotient work inside the loop.
Source hypothesis
The existing direct matcher cannot see this form. Range and bit-field relations can show an equivalent index.
Meaning conditions
The transform must keep cast and integer-wrap semantics. Signed correction needs a separate proof.
Open boundary
The first missed fold remains unknown. Wide loads and stores already exist.
31: address work 0260-0390, backedge 0480 to 0260. LDG128: 03a0. STG128: 0470.
Q-E | Nested output-address alignment
7B: 28 | LLVM KnownBits and alignment analysis
Figure 11. Q-E: One copy has different load and store widths. This figure is a schematic, not a timing chart.Figure 12. Q-E: Static path toward the extent facts. This figure is a schematic, not a timing chart.
Computation
The GQA copy repeats four KV heads to make 28 heads. Each thread copies four half values.
Observed code
The input uses one 64-bit load. The output uses four 16-bit stores.
Source hypothesis
The output query reaches depth 6 before the extent facts. This depth path can explain the alignment difference.
Meaning conditions
The extent is a multiple of 128. The byte offset is a multiple of eight.
Open boundary
The actual query stack has no trace. GQA materialization removal would be a separate producer change.
7B: 2, 4 | Possible LLVM NVPTX source implementation
Figure 13. Q-F: The denominator work becomes visible late. This figure is a schematic, not a timing chart.Figure 14. Q-F: A possible separation of exact division. This figure is a schematic, not a timing chart.
Computation
Index and mask loops divide different numerators by an invariant denominator. The original arithmetic has 64-bit semantics.
Observed code
ptxas adds reciprocal setup and quotient correction. The setup occurs inside the observed loop.
Source hypothesis
A NVPTX source implementation can expose denominator work before PTX emission. LLVM can then share that work.
Meaning conditions
Each numerator needs exact quotient correction. The implementation must keep the 64-bit fallback and signed corner cases.
Open boundary
This investigation has no exact algorithm or source patch. This observation does not establish a defect in LLVM.
Source and code evidence
NVPTX division emission / ptxas expansion is outside upstream scope
No exact source implementation or algorithm is prepared.
Dispatch 2 uses a 32-bit fast path and a separate 64-bit fallback.
2: reciprocal setup and correction 02e0-0410. Backedge: 05c0 to 0280.
Q-G | One denominator for exact Softmax division
7B: 33 | Possible LLVM NVPTX source implementation
Figure 15. Q-G: Softmax has three scans. This figure is a schematic, not a timing chart.Figure 16. Q-G: Eight quotients use the same denominator. This figure is a schematic, not a timing chart.
Computation
Softmax uses maximum, sum and output scans. Eight quotients use one row denominator.
Observed code
Eight PTX divisions use the same denominator. Offline SASS has eight reciprocal operations on R39.
Source hypothesis
A source implementation can share denominator setup. Each quotient must keep its own rounding and error correction.
Meaning conditions
Reciprocal multiplication can change exact division results. Zero, subnormal, NaN and infinity cases need the original behavior.
Open boundary
The exp recomputation boundary remains unknown. The frozen common Linalg decomposition shares one exp tensor.
7B examples: 1, 3, 7 | Common SCF, arithmetic or LLVM loop analysis / boundary open
Figure 17. Q-H: The small loop still has a backedge. This figure is a schematic, not a timing chart.Figure 18. Q-H: The missing range proof. This figure is a schematic, not a timing chart.
Computation
The tile loop has step 32. Its upper bound is the smaller of remaining and 32.
Observed code
The backedge remains in O3 and SASS. Signed truncation prevents an unconditional zero-or-one iteration claim.
Source hypothesis
A valid range proof can remove the backedge. The responsible common transform remains unknown.
Meaning conditions
The proof must follow the signed range through i64-to-i32 truncation. It must also cover the outer loop bounds.
Open boundary
The SCF-to-CF snapshot is missing. A producer-only loop change would not qualify.
Source and code evidence
SCF-to-CF boundary and common loop analysis: not localized