Quick keys
— Navigate with ← → or spacebar.
F
fullscreen ·
S
speaker view ·
O
overview ·
Alt+Click
zoom
Navigation
×
<style> :root { --q-ink:#142d48; --q-muted:#516477; --q-line:#ccd6dd; --q-card:#fff; --q-blue:#17639a; --q-green:#16715f; --q-orange:#a74919; --q-tint:#edf3f6; } html.dark { --q-ink:#edf3f8; --q-muted:#b8c7d5; --q-line:#496074; --q-card:#1c2e3e; --q-blue:#8bcef5; --q-green:#89dbbf; --q-orange:#ffc197; --q-tint:#20384b; } .reveal .slides section .qslide { box-sizing:border-box; width:100%; height:820px; padding:24px 44px 22px; color:var(--q-ink); font-family:Inter,Arial,sans-serif; display:flex; flex-direction:column; text-align:left; font-size:29px; line-height:1.34; } .reveal .slides section .qslide h1 { color:var(--q-ink); font-size:72px; font-weight:700; line-height:1.09; letter-spacing:-1.3px; text-transform:none; margin:25px 0 22px; max-width:1340px; } .reveal .slides section .qslide h2 { color:var(--q-ink); font-size:51px; font-weight:650; line-height:1.14; letter-spacing:-.7px; text-transform:none; margin:10px 0 30px; } .reveal .slides section .qslide h3 { color:var(--q-ink); font-size:30px; font-weight:650; line-height:1.25; text-transform:none; margin:0 0 16px; } .reveal .slides section .qslide p { line-height:1.38; margin:0 0 20px; font-size:inherit; } .reveal .slides section .qslide .lead { font-size:38px; line-height:1.3; } .reveal .slides section .qslide .muted { color:var(--q-muted); } .reveal .slides section .qslide .small { font-size:24px; } .reveal .slides section .qslide .tiny { font-size:21px; } .reveal .slides section .qslide .body { flex:1 0 auto; min-height:0; } .reveal .slides section .qslide .grid { display:grid; grid-template-columns:minmax(0,1fr) minmax(0,1fr); gap:32px; align-items:start; } .reveal .slides section .qslide .grid.wide { grid-template-columns:minmax(0,1.2fr) minmax(0,1fr); } .reveal .slides section .qslide .card { padding:25px 28px; border:1px solid var(--q-line); border-radius:10px; background:var(--q-card); box-shadow:none; } .reveal .slides section .qslide .card p:last-child { margin-bottom:0; } .reveal .slides section .qslide .blue { border-top:5px solid var(--q-blue); } .reveal .slides section .qslide .green { border-top:5px solid var(--q-green); } .reveal .slides section .qslide .orange { border-top:5px solid var(--q-orange); } .reveal .slides section .qslide .callout { margin-top:24px; padding:19px 25px; background:var(--q-tint); border-left:5px solid var(--q-blue); font-size:29px; line-height:1.33; } .reveal .slides section .qslide.eval .callout { border-left-color:var(--q-green); } .reveal .slides section .qslide.future .callout { border-left-color:var(--q-orange); } .reveal .slides section .qslide footer { font-size:19px; line-height:1.3; color:var(--q-muted); border-top:1px solid var(--q-line); padding-top:11px; margin-top:18px; } .reveal .slides section .qslide table { width:100%; margin:0; font-size:27px; line-height:1.3; border-collapse:collapse; } .reveal .slides section .qslide th { font-weight:600; color:var(--q-muted); font-size:22px; border-bottom:2px solid var(--q-line); padding:12px 15px; } .reveal .slides section .qslide td { padding:18px 15px; border-bottom:1px solid var(--q-line); vertical-align:top; } .reveal .slides section .qslide table.compact { font-size:24px; } .reveal .slides section .qslide table.compact td { padding:12px 14px; } .reveal .slides section .qslide .equation { border-radius:10px; padding:20px; background:var(--q-tint); font-size:34px; text-align:center; margin:14px 0 26px; } .reveal .slides section .qslide .katex-display, .reveal .slides section .qslide .MathJax_Display { margin:.35em 0; } .reveal .slides section .qslide img.chart { box-sizing:border-box; max-width:100%; max-height:425px; width:100%; background:#fff; border-radius:10px; box-shadow:none; padding:14px; object-fit:contain; } .reveal .slides section .qslide .metric { color:var(--q-blue); font-size:66px; line-height:1.15; font-weight:700; margin:4px 0 16px; } .reveal .slides section .qslide .flow { display:grid; grid-template-columns:1fr 40px 1fr 40px 1fr; gap:14px; align-items:center; margin:40px 0; } .reveal .slides section .qslide .flow .arrow { font-size:42px; color:var(--q-muted); text-align:center; } .reveal .slides section .qslide .flow .card { min-height:160px; } .reveal .slides section .qslide .cover-meta { margin-top:32px; font-size:24px; line-height:1.5; } .reveal .slides section .qslide a { color:var(--q-blue); text-decoration:underline; text-underline-offset:3px; } .reveal .slides section .qslide a:focus-visible { outline:3px solid var(--q-orange); outline-offset:5px; } .reveal .slides section .qslide > * { flex-shrink:0; } .reveal .slides section .qslide.cover { justify-content:center; } .reveal .slides section .qslide.cover h1 { margin:0 0 38px; } .reveal .slides section .qslide.cover .author { font-size:38px; margin:0; } .reveal .slides section .qslide .ack-photos { display:grid; grid-template-columns:minmax(0,1fr) minmax(0,1fr); gap:24px; margin:0 0 26px; } .reveal .slides section .qslide .ack-photos img { width:100%; height:460px; object-fit:contain; border-radius:0; box-shadow:none; background:transparent; margin:0; } .reveal .slides section .qslide.thanks { align-items:center; justify-content:center; text-align:center; } .reveal .slides section .qslide.thanks h1 { font-size:100px; margin:0 0 40px; } .reveal .slides section .qslide.thanks .questions { font-size:44px; margin-bottom:60px; } .reveal .slides section .qslide.thanks p { font-size:28px; } .reveal .slides section .qslide .tail-copy .equation { padding:12px; margin:0 0 18px; font-size:29px; } .reveal .slides section .qslide .tail-copy p { margin-bottom:16px; } @media print { .reveal .slides section .qslide { color:#142d48; --q-ink:#142d48; --q-muted:#516477; --q-line:#ccd6dd; --q-card:#fff; --q-blue:#17639a; --q-green:#16715f; --q-orange:#a74919; --q-tint:#edf3f6; } } </style> <div class="qslide cover"> <h1>Reliable LLM inference<br>under resource constraints</h1> <p class="author">Mohsen Hariri</p> <p class="cover-meta">Advisor: Prof. Vipin Chaudhary · Committee: Prof. Jing Ma, Prof. Sanmukh Kuppannagari</p> </div> Note: Time: 0:20. Cumulative: 0:20. I study whether reducing a model's memory cost lets it return correct answers more often at a fixed resource budget. Compression determines which inference procedures fit. Evaluation measures how often those procedures return a correct answer. --- <div class="qslide"> <h2>Can lower inference cost improve answer accuracy?</h2> <div class="body"> <p class="lead">Compare the full system at a fixed memory and compute budget.</p> <div class="flow"> <div class="card blue"><h3>Compression</h3><p class="small">Weights and KV cache determine what fits.</p></div> <div class="arrow" aria-hidden="true">→</div> <div class="card"><h3>Inference procedure</h3><p class="small">Spend capacity on longer or repeated attempts.</p></div> <div class="arrow" aria-hidden="true">→</div> <div class="card green"><h3>Evaluation</h3><p class="small">Measure the answer delivered to the user.</p></div> </div> <p class="callout">More candidates help only if the selector returns a better answer.</p> </div> </div> Note: Time: 0:50. Cumulative: 1:10. A smaller representation can let a model fit or leave room for more candidates. The comparison has two steps: measure how compression changes the candidates, then measure what happens when we spend the savings on more inference. The selector is part of the system in both cases. --- <div class="qslide"> <h2>Memory use and answer accuracy</h2> <div class="body"> <div class="equation">$$M_{\mathrm{KV}}\simeq LBT H_{\mathrm{KV}}d_h(b_K+b_V)/8$$</div> <div class="grid"> <div class="card blue"><h3>Memory use</h3><p>Weights + cache + activations + temporary buffers.</p><p class="small muted">More concurrent attempts increase B.<br>Longer reasoning increases T.</p></div> <div class="card green"><h3>Answer reliability</h3><p>Numerical changes can alter the answer distribution.</p><p class="small muted">Finite response banks leave uncertainty about that distribution.</p></div> </div> <p class="callout">Compare peak memory and total GPU time. Also report latency, tokens, and verifier cost.</p> </div> <footer>L: layers; B: concurrent sequences; T: cached tokens; H<sub>KV</sub>: KV heads; d<sub>h</sub>: head dimension; b: bits. Payload formula excludes overhead.</footer> </div> Note: Time: 1:00. Cumulative: 2:10. The formula counts cache payload. Scales, residual buffers, padding, and decompression workspace add to peak memory. Longer sequences increase T; concurrent attempts increase B. Lower precision reduces both costs but can change generated tokens. Repeating a question adds evidence about that question without adding an independent task. --- <div class="qslide"> <h2>Weight and cache compression</h2> <div class="body"> <table><thead><tr><th>Family</th><th>Mechanism</th><th>Tradeoff to test</th></tr></thead><tbody> <tr><td>Weight quantization<br><span class="small muted">GPTQ; AWQ</span></td><td>Change weight precision using calibration information.</td><td>More compression; accuracy can change.</td></tr> <tr><td>Exact weight coding<br><span class="small muted">NeuZip coding; DFloat11</span></td><td>Encode redundancy; reconstruct numerical values.</td><td>Preserve values; pay for decoding.</td></tr> <tr><td>KV quantization<br><span class="small muted">KIVI; mixed K/V bits</span></td><td>Compress activations as the cache grows.</td><td>Save cache space; perturb attention.</td></tr> <tr><td>Cache management<br><span class="small muted">PagedAttention</span></td><td>Reduce allocation waste and share prefixes.</td><td>Complementary to precision changes.</td></tr> </tbody></table> </div> <footer>Frantar et al., ICLR 2023; Lin et al., MLSys 2024; Hao et al., 2024; Liu et al., ICML 2024; Kwon et al., SOSP 2023.</footer> </div> Note: Time: 1:15. Cumulative: 3:25. GPTQ compensates for weight error using approximate second-order information. AWQ chooses scales using activations. Entropy coding preserves values and adds decoding work; NeuZip also has near-lossless settings. KIVI changes grouping dimensions, while mixed K/V precision changes bit widths. PagedAttention reduces allocation waste and shares cached prefixes. --- <div class="qslide"> <h2>DFloat11 trades decoding work for GPU residency</h2> <div class="body grid wide"> <div><img class="chart" src="/assets/slides/2026-09-30-phd-qualifying-examination/dfloat-size.svg" alt="Llama 3.3 70B stored model size: BF16 141.11 GB; DFloat11 95.40 GB."><p class="small muted">Llama 3.3 70B Instruct · Table 1</p></div> <div class="card blue"><p class="metric">67.61%</p><p>of the original stored model size.</p><p>BF16 weights recover exactly.</p><p class="small muted">Low-entropy exponents + hierarchical lookup tables + block-level GPU decompression.</p></div> </div> <p class="callout">Reported throughput gains compare against CPU-offloaded BF16. Resident BF16 still avoids decompression work.</p> <footer>Zhang, Hariri et al., “70% Size, 100% Accuracy,” NeurIPS 2025. Table 1, §3, Appendix K. Stored size is not peak process memory.</footer> </div> Note: Time: 1:10. Cumulative: 4:35. BF16 allocates eight bits to exponents, although the studied weights have exponent entropy near 2.6 bits. Coding those exponents while retaining the sign and fraction gives roughly eleven bits per weight. The throughput comparisons use CPU-offloaded BF16. When BF16 already fits on the GPU, the comparison must account for decompression overhead. --- <div class="qslide"> <h2>The same cache payload can give different accuracy</h2> <div class="body grid wide"> <div><img class="chart" src="/assets/slides/2026-09-30-phd-qualifying-examination/kv-allocation.svg" alt="GSM8K accuracy: K2/V4 54.7 percent; K4/V2 75.2 percent; K4/V4 75.4 percent."><p class="small muted">Llama 3.1 8B Instruct · GSM8K · one-shot</p></div> <div class="card blue"><h3>Keys change attention weights</h3><p>Values change the content being combined.</p><p class="small">K2/V4 and K4/V2 both use 6 payload bits per K/V pair.</p><p class="small">K4/V2 uses 25% fewer payload bits than K4/V4.</p></div> </div> <p class="callout">Test which layers and heads benefit from higher key precision.</p> <footer>Hariri et al., “Quantize What Counts,” Findings of ACL 2026, Table 3. Token-wise Optimum Quanto; equal K/V sizes; overhead excluded.</footer> </div> Note: Time: 1:20. Cumulative: 5:55. At the same cache payload, switching from K2/V4 to K4/V2 changes reported accuracy from 54.7 to 75.2 percent. K4/V4 reaches 75.4 percent with more bits. Those are point estimates; the table has no equivalence test. I would test whether attention sensitivity predicts where an additional bit improves accuracy across contexts and backends. --- <div class="qslide"> <h2>From attention error to answer accuracy</h2> <div class="body"> <div class="equation">$$\delta o\approx J_{\mathrm{softmax}}\!\left(q\delta K^\top/\sqrt{d_h}\right)V+A\delta V$$</div> <div class="grid"> <div class="card"><h3>Attention error</h3><p>Key and value errors enter attention through different paths.</p><p class="small muted">Allocation can depend on the query, layer, and attention distribution.</p></div> <div class="card"><h3>Answer accuracy</h3><p>Whether the error surrogate predicts delivered-answer accuracy.</p><p class="small muted">Measure whether additional attempts offset numerical loss and runtime cost.</p></div> </div> <p class="callout">Test whether spending saved memory on more attempts improves the selected answer.</p> </div> <footer>First-order attention perturbation; higher-order terms omitted. Compare the approximation with measured attention error and task accuracy.</footer> </div> Note: Time: 1:00. Cumulative: 6:55. Key error passes through the softmax and values; attention averages value error. A tensor norm alone misses those directions. I would measure attention-output error and answer accuracy on held-out prompts, then test whether the extra attempts made possible by compression offset any accuracy loss. --- <div class="qslide eval"> <h2>Where extra inference compute goes</h2> <div class="body"> <table><thead><tr><th>Procedure</th><th>Potential benefit</th><th>Failure or cost to measure</th></tr></thead><tbody> <tr><td>One longer trajectory<br><span class="small muted">Chain-of-thought</span></td><td>More intermediate reasoning</td><td>Longer cache; an early mistake can persist</td></tr> <tr><td>Completed candidates<br><span class="small muted">Self-consistency; verification</span></td><td>Several attempts, then select</td><td>Correlated wrong answers; selector errors and cost</td></tr> <tr><td>Search over prefixes<br><span class="small muted">Tree of Thoughts</span></td><td>Revisit intermediate decisions</td><td>Branch state; discarded work; evaluator bias</td></tr> </tbody></table> <p class="callout">The full generation and selection procedure defines the system being evaluated.</p> </div> <footer>Wei et al., NeurIPS 2022; Wang et al., ICLR 2023; Yao et al., NeurIPS 2023; Lightman et al., ICLR 2024; Snell et al., ICLR 2025.</footer> </div> Note: Time: 1:20. Cumulative: 8:15. Completed attempts can form an independent response bank. Search leaves share prefixes and decisions, so they need a different statistical treatment. Self-consistency selects by agreement; a verifier predicts correctness. Both incur costs and can select a wrong answer. Snell and colleagues show that compute allocation depends on the problem, motivating cost-quality curves. --- <div class="qslide eval"> <h2>Finding a correct candidate differs from returning one</h2> <div class="body"> <table><thead><tr><th>Quantity</th><th>Question</th><th>Per-task probability under iid attempts</th></tr></thead><tbody> <tr><td>Single-attempt accuracy</td><td>Does one attempt succeed?</td><td>$$p_i$$</td></tr> <tr><td>Discovery: Pass@k</td><td>Is any candidate correct?</td><td>$$1-(1-p_i)^k$$</td></tr> <tr><td>All-attempt success</td><td>Does every attempt succeed?</td><td>$$p_i^k$$</td></tr> <tr><td>Delivered-answer accuracy</td><td>Does the selector return a correct answer?</td><td>Measure the specified selector</td></tr> </tbody></table> </div> <footer>Chen et al., “Evaluating Large Language Models Trained on Code,” 2021. A selector restricted to the same bank is bounded by oracle discovery.</footer> </div> Note: Time: 1:10. Cumulative: 9:25. Pass@k measures whether the bank contains a correct answer. Majority vote or a learned verifier may still choose a wrong one. Discovery bounds selected-answer accuracy when the selector must return a member of that same bank. Repairing or generating another answer changes the system. The formulas assume iid attempts for a fixed task and configuration. --- <div class="qslide eval"> <h2>Discovery and repeatability can rank models differently</h2> <div class="body"> <p class="small">A: success probability 0.6 on every task. B: always correct on half the tasks, always wrong on half.</p> <img class="chart" style="max-height:340px;" src="/assets/slides/2026-09-30-phd-qualifying-examination/discovery-repeatability.svg" alt="Synthetic systems: A has higher discovery as k grows; B has higher all-attempt success from k equals 2 onward."> <div class="grid" style="margin-top:12px;"><div class="card blue small"><h3>At k = 4 · system A</h3><p>Discovery: 0.9744 · All correct: 0.1296</p></div><div class="card orange small"><h3>At k = 4 · system B</h3><p>Discovery: 0.5 · All correct: 0.5</p></div></div> </div> <footer>Analytic example with equally weighted tasks.</footer> </div> Note: Time: 1:10. Cumulative: 10:35. System A has mean accuracy 0.6, compared with 0.5 for B. At four attempts, A has discovery 0.9744 but all-attempt success only 0.1296. B stays at 0.5 for both because its success is concentrated on half the tasks. The ranking changes because the objectives differ; there is no estimation noise in this example. --- <div class="qslide eval"> <h2>Bayesian uncertainty from repeated responses</h2> <div class="body"> <div class="equation">$$p_i\mid D\sim\operatorname{Beta}(a+c_i,\ b+N-c_i)$$</div> <div class="grid"> <div class="card green"><h3>Uncertainty</h3><p>Estimate uncertainty for the chosen outcome rubric.</p><p class="small muted">Compare the posterior difference against a prespecified tolerance.</p></div> <div class="card"><h3>Point estimates</h3><p>With a uniform prior and equal N:</p><p>$$\overline{\mathbb E[p_i\mid D]}=\frac{N}{N+2}\operatorname{Avg}@N+\frac1{N+2}$$</p><p class="small muted">The point-estimate ordering is unchanged.</p></div> </div> <p class="callout">Uncertainty over attempts on fixed questions differs from uncertainty over new questions.</p> </div> <footer>Hariri, Samandar, Hinczewski & Chaudhary, “Don't Pass@k,” ICLR 2026. Binary case shown; the published framework also uses categorical outcomes.</footer> </div> Note: Time: 1:20. Cumulative: 11:55. With a uniform prior and equal trial counts, posterior mean accuracy preserves the average-accuracy ranking. The posterior adds uncertainty estimates and extends to categorical rubrics. I would compare posterior differences between systems. Fixed-benchmark uncertainty also leaves uncertainty about new questions, which needs paired question resampling and new response banks. --- <div class="qslide eval"> <h2>Ranking methods encode different assumptions</h2> <div class="body"> <table><thead><tr><th>Family</th><th>What the ordering summarizes</th><th>Assumption or information loss</th></tr></thead><tbody> <tr><td>Direct scores</td><td>Mean performance under a rubric</td><td>Task weights and chosen utility</td></tr> <tr><td>Bradley–Terry</td><td>Pairwise wins</td><td>A latent pairwise model; score magnitude can be lost</td></tr> <tr><td>Rasch / item response</td><td>Ability relative to task difficulty</td><td>Latent structure must fit the tasks</td></tr> </tbody></table> <div class="callout">20 models × 4 math benchmarks, up to 80 trials.<br>The reference ranking uses full-budget Bayes<sub>U</sub>@80.</div> <p class="small muted" style="margin-top:20px;">Scorio.jl compares these methods on the same response tensor.</p> </div> <footer>Bradley & Terry, 1952; Rasch, 1960; Hariri et al., “Ranking Reasoning LLMs,” ACL 2026; “Scorio.jl,” JuliaCon.</footer> </div> Note: Time: 1:05. Cumulative: 13:00. The ranking study uses the same response tensor across methods. It compares rankings with full-budget Bayes_U@80 and with each method's own full-budget ordering. Differences can come from the target or from sampling uncertainty. Greedy responses used as an empirical prior can reduce variance while biasing rankings when greedy and sampled behavior differ. Scorio.jl implements these comparisons. --- <div class="qslide future"> <h2>Next experiments</h2> <div class="body"> <table><thead><tr><th>Limitation</th><th>Direction</th><th>Evidence needed</th></tr></thead><tbody> <tr><td>Compression changes memory use and numerical error.</td><td>Spend compression savings under a fixed budget.</td><td>Delivered-answer gains after all costs.</td></tr> <tr><td>A scalar hides repeated-success behavior.</td><td>TailPass@k; compare Geom@k.</td><td>Higher selected-answer accuracy on new tasks and response banks.</td></tr> <tr><td>Selection can fail on new tasks or models.</td><td>Runtime signals; semantic decoding; paired tasks.</td><td>Selection accuracy on held-out models and domains.</td></tr> </tbody></table> </div> <footer>Hariri et al., “Geom@k,” “Success Has a Shape,” “Runtime Signals,” and “When Math Becomes Physics”; Anonymized, “Semantic Sampling.”</footer> </div> Note: Time: 0:45. Cumulative: 13:45. I would first test whether compression savings improve delivered answers. Success profiles may help choose a system, and runtime signals or semantic decoding may improve its answer selection. Each comparison needs new tasks and response banks so that the same data does not determine both the choice and its measured benefit. --- <div class="qslide future"> <h2>Compression at fixed attempts and fixed resources</h2> <div class="body"> <div class="grid"> <div class="card blue"><h3>Fixed attempts</h3><p>Hold the checkpoint, prompts, decoder, selector, and attempt count fixed.</p><p class="small muted">Compare BF16, DFloat11, and supported low-bit weights; vary KV allocation.</p></div> <div class="card orange"><h3>Fixed resources</h3><p>Hold peak-memory and GPU-time limits fixed.</p><p class="small muted">Let savings fund more candidates or longer reasoning; charge all selection cost.</p></div> </div> <p class="callout">I expect the largest gains when BF16 weights barely exceed memory or the KV cache dominates it.</p> <p style="margin-top:24px;">The gain may disappear if decompression is too slow or numerical error lowers answer quality.</p> </div> <footer>Mark configurations that exceed memory as infeasible. Compare CPU-offloaded BF16 separately.</footer> </div> Note: Time: 1:25. Cumulative: 15:10. Fixed attempts isolate the representation change. Fixed resources test the additional inference bought by the savings. A BF16 model that exceeds the memory cap is infeasible; offloading is a separate baseline with its own cost. For cache precision, I would compare fixed key-favored, norm-based, and sensitivity-based allocations. Discovery helps distinguish candidate failures from selection failures. --- <div class="qslide future"> <h2>TailPass@k: success at every threshold</h2> <div class="body grid wide"> <div><img class="chart" style="max-height:365px;" src="/assets/slides/2026-09-30-phd-qualifying-examination/tail-profile.svg" alt="Synthetic k equals 4 tail profiles: system A decreases from 0.9744 to 0.1296; system B remains at 0.5."><p class="small muted">Synthetic example, k = 4.</p></div> <div class="tail-copy"><div class="equation">$$S_{k,t}=\frac1M\sum_i\Pr(X_{i,k}\geq t)$$</div><p class="small">t = 1: at least one correct<br>t = k: all attempts correct</p><p class="small">Geom@k takes the geometric mean of the two dataset-average endpoints.</p></div> </div> <p class="callout">Compare system choices based on accuracy, Pass@k, Geom@k, and the profile using new response banks.</p> <footer>Hariri et al., “Success Has a Shape” and “Geom@k.” Analytic example.</footer> </div> Note: Time: 1:20. Cumulative: 16:30. TailPass@k records the probability of reaching each success count. Under iid attempts, the average across thresholds equals mean accuracy. Geom@k combines dataset-average discovery and all-attempt success with a geometric mean. I would compare the system choices these scores produce against held-out selected-answer accuracy. Finite-bank subsets and predictions for a future bank answer different questions. --- <div class="qslide future"> <h2>Runtime signals and semantic decoding</h2> <div class="body grid"> <div class="card orange"><h3>Runtime signals</h3><p>Use completion status, length, formatting, and token-confidence summaries to select candidates.</p><p class="small muted">Baselines: majority vote, length only, random selection, learned verifier.</p><p class="small">Hold out whole problems, then models or domains. Charge predictor cost.</p></div> <div class="card orange"><h3>Semantic Sampling for math</h3><p>Rescore candidates using probability mass in embedding neighborhoods.</p><p class="small muted">Baselines: greedy, nucleus, matched filters, random neighborhoods.</p><p class="small">Test meaning-changing neighbors and lookup overhead. Determinism alone does not establish correctness.</p></div> </div> <p class="callout">Primary outcome: delivered-answer accuracy at matched total cost.</p> <footer>Hariri et al., “Runtime Signals”; Anonymized, “Semantic Sampling.” Medical adaptation: Yu, Hariri et al., MICCAI 2026 (accepted).</footer> </div> Note: Time: 1:10. Cumulative: 17:40. A predictor can separate correct and incorrect traces in aggregate yet fail to select a correct answer for one problem. I would measure selection accuracy alongside AUROC. Whole problems stay together in data splits, and runtime features exclude gold answers. For semantic decoding, nearby embeddings can have opposite meanings. The mathematics comparison must include those cases and the cost of neighborhood lookup. --- <div class="qslide future"> <h2>Transfer to physics and agent tasks</h2> <div class="body grid"> <div class="card orange"><h3>Math2Physics</h3><p>Pair formulations that share a central derivation.</p><p class="small">Test whether extra computation closes the delivered-answer gap.</p><p class="small muted">Preserve pairs in analysis; review ambiguity and added interpretation.</p></div> <div class="card orange"><h3>BenignPass@k</h3><p>Require success and a separate benignness condition.</p><p class="small">$$J_{i,k,s,b}=\Pr(X_{i,k}\geq s,\ Z_{i,k}\geq b)$$</p><p class="small muted">For s = 1, b = k: some success and no harmful attempt.</p></div> </div> <p class="callout">A selected correct answer does not describe errors or harm elsewhere in the attempted work.</p> <footer>Hariri et al., “When Math Becomes Physics”; Anonymized, “BenignPass@k.”</footer> </div> Note: Time: 1:10. Cumulative: 18:50. Math2Physics tests whether the same derivation becomes harder when expressed as a physical problem. Extra attempts may increase discovery without improving the selected answer. BenignPass measures success and benignness jointly. For thresholds below k, different attempts can satisfy the two counts. Those thresholds must specify the behavior required of the agent. --- <div class="qslide eval"> <h2>Compression gains measured in answer accuracy</h2> <div class="body"> <div class="grid"> <div class="card blue"><h3>Compression</h3><p>Measure memory savings alongside decompression cost and numerical error.</p></div> <div class="card green"><h3>Evaluation</h3><p>Measure the selected answer and uncertainty across tasks and attempts.</p></div> </div> <p class="lead" style="margin-top:45px;">I propose to test when spending compression savings on additional inference improves delivered-answer accuracy at a fixed resource budget.</p> <p class="callout">If accuracy does not improve, measure whether numerical error, runtime cost, or selection prevents the gain.</p> </div> </div> Note: Time: 0:40. Cumulative: 19:30. I would test whether additional inference repays the numerical and runtime costs of compression. The result should identify memory limits and workloads where answer accuracy improves, and where decompression, candidate quality, or answer selection prevents a gain. --- <div class="qslide acknowledgements"> <h2>Acknowledgements</h2> <div class="body"> <div class="ack-photos"> <img src="/assets/slides/2026-09-30-phd-qualifying-examination/research-group.webp" alt="Vipin Chaudhary research group"> <img src="/assets/slides/2026-09-30-phd-qualifying-examination/group-dinner.webp" alt="Group dinner"> </div> <p class="small">Thank you to Prof. Vipin Chaudhary, my committee, collaborators, and research group.</p> </div> <footer>NSF awards 2117439, 2112606, 2320952 · High Performance Computing Center at CWRU</footer> </div> Note: Time: 0:20. Cumulative: 19:50. Thank you to my advisor, committee, collaborators, and research group, and to NSF and the CWRU HPC Center for their support. --- <div class="qslide thanks"> <h1>Thank you!</h1> <p class="questions">Questions?</p> <p><a href="mailto:mohsen.hariri@case.edu">mohsen.hariri@case.edu</a></p> <p><a href="https://mohsenhariri.github.io">mohsenhariri.github.io</a></p> </div> Note: Time: 0:10. Cumulative: 20:00. Thank you. --- <div class="qslide future"> <h2>Allocating bits by attention sensitivity</h2> <div class="body"> <div class="equation">$$D=c_K2^{-2b_K}+c_V2^{-2b_V},\qquad b_K+b_V=b_{\mathrm{tot}}$$</div> <div class="equation">$$b_K-b_V=\tfrac12\log_2(c_K/c_V)$$</div> <p>The expression is an unconstrained continuous optimum under the stated error model.</p> <p>Supported integer precisions, clipping, metadata, and kernel costs constrain an implementation.</p> <p class="callout">Compare sensitivity estimates with fixed, swapped, and norm-based allocations on held-out prompts.</p> </div> <footer>Hariri et al., “Quantize What Counts,” Findings of ACL 2026. Compare predicted error with task accuracy.</footer> </div> Note: Substitute bV = btot - bK and differentiate. The stationary point equates the two weighted error contributions. Reversing their ratio reverses the preferred bit split; supported precisions and boundary constraints limit the choice. Estimate the coefficients on calibration data, then check both attention error and answer accuracy on held-out tasks. --- <div class="qslide eval"> <h2>Categorical uncertainty</h2> <div class="body"> <div class="equation" style="font-size:28px; padding:12px;">$$\boldsymbol\theta_i\mid D\sim\operatorname{Dirichlet}(\boldsymbol\nu_i),\qquad \boldsymbol\nu_i=\boldsymbol\alpha_i+\mathbf n_i$$</div> <p>Let a<sub>i</sub> = ∑<sub>j</sub> ν<sub>ij</sub> and μ<sub>i</sub> = ∑<sub>j</sub> w<sub>j</sub>ν<sub>ij</sub>/a<sub>i</sub>.</p> <div class="equation" style="font-size:28px; padding:12px;">$$\operatorname{Var}(\mathbf w^\top\boldsymbol\theta_i\mid D)=\frac{\sum_jw_j^2\nu_{ij}/a_i-\mu_i^2}{a_i+1}$$</div> <p class="small">Independent task posteriors give variance M<sup>−2</sup>∑<sub>i</sub>Var(w<sup>⊤</sup>θ<sub>i</sub> | D) for the fixed-benchmark average.</p> <p class="callout small">Fix the rubric before evaluation. Check prior sensitivity and model assumptions.</p> </div> <footer>Hariri et al., “Don't Pass@k,” ICLR 2026. Posterior draws can compare differences without relying on a Gaussian approximation.</footer> </div> Note: The moments assume a fixed rubric and independent task parameters. A hierarchical model would require a different joint calculation. Malformed or truncated responses remain outcomes. Fixing the rubric before inspecting results prevents choosing weights to favor a particular ranking. --- <div class="qslide eval"> <h2>Observed subsets and future attempts</h2> <div class="body"> <p>For N recorded attempts, c successes, and k ≤ N:</p> <div class="equation" style="font-size:28px; padding:12px;">$$\widehat S_{k,t}=\sum_{j=t}^{k}\frac{\binom cj\binom{N-c}{k-j}}{\binom Nk}$$</div> <p>With p | D ∼ Beta(a, b), the posterior prediction for a future iid bank is:</p> <div class="equation" style="font-size:28px; padding:12px;">$$\Pr(X_k\geq t\mid D)=\sum_{j=t}^{k}\binom kj\frac{B(a+j,b+k-j)}{B(a,b)}$$</div> <p class="callout small">k > N requires model-based extrapolation. Search leaves need a dependency model.</p> </div> <footer>Chen et al., 2021; Hariri et al., “Success Has a Shape.” Impossible binomial terms are zero.</footer> </div> Note: The first expression counts subsets of the observed bank and reduces to the combinatorial Pass@k estimator at t = 1. The second integrates a future-binomial probability over the posterior for p. Here a and b are posterior parameters. Uncertainty about the event probability differs from variation in the next realized success count. --- <div class="qslide"> <h2>So far ...</h2> <div class="body"> <table class="compact"><thead><tr><th>Work</th><th>Venue</th><th>Role</th></tr></thead><tbody> <tr><td>DFloat11</td><td>NeurIPS 2025</td><td>Exact weight compression</td></tr> <tr><td>Quantize What Counts</td><td>Findings of ACL 2026</td><td>Mixed K/V precision</td></tr> <tr><td>Don't Pass@k</td><td>ICLR 2026</td><td>Bayesian evaluation</td></tr> <tr><td>Ranking Reasoning LLMs</td><td>ACL 2026</td><td>Ranking assumptions</td></tr> <tr><td>Scorio.jl</td><td>JuliaCon</td><td>Shared ranking interface</td></tr> <tr><td>Medical Image Spatial Grounding<br>with Semantic Sampling</td><td>MICCAI 2026 · accepted</td><td>Grounding application and evaluation limits</td></tr> </tbody></table> </div> </div> Note: The published work covers weight and cache compression, Bayesian evaluation, ranking, and the Scorio.jl implementation. The accepted medical study applies semantic sampling to grounding. It reports 33,864 questions, mostly CT, and excludes malformed answers from its spatial-correctness measure. A follow-up would also report correctness over all attempts. --- <div class="qslide future"> <h2>Next</h2> <div class="body"> <table class="compact"><thead><tr><th>Work</th><th>Question to investigate</th></tr></thead><tbody> <tr><td>Geom@k</td><td>When does an endpoint balance help decisions?</td></tr> <tr><td>Test-Time Scaling survey</td><td>How should protocols and budgets be compared?</td></tr> <tr><td>Runtime Signals</td><td>Can cheap features improve held-out selection?</td></tr> <tr><td>TailPass@k</td><td>Which success thresholds matter operationally?</td></tr> <tr><td>Math2Physics</td><td>Does reliability transfer across formulations?</td></tr> <tr><td>Decoding survey</td><td>Which operator choices change generation?</td></tr> <tr><td>BenignPass@k</td><td>How should success and harm be measured jointly?</td></tr> <tr><td>Semantic Sampling for math</td><td>Does semantic rescoring improve reasoning?</td></tr> </tbody></table> </div> </div> Note: I would compare these methods by their effect on held-out answer accuracy and cost. The surveys organize the inference procedures and decoding operators. The remaining studies test repeated success, answer selection, transfer, and joint success and harm.