Hi everyone,
I’m an independent researcher from India who just completed empirical work on decision-layer architectures for LLM coding agents entirely on a single GPU.
Key result: A DPO-trained clarification judge (QLoRA on Qwen2.5-Coder-7B) raises one-shot pass@1 from 0.419 → 0.639 on a frozen 14B generator — +22 percentage points, McNemar p<0.000001, seed-stable across 3 runs.
Paper is ready to submit to arXiv but I need one endorsement from someone who has submitted 3+ papers to cs.AI on arXiv.
If you can help, one click here endorses me:
Or email endorse@arxiv.org with code GIH67R
Would really appreciate any help from this community! ![]()