Phase 5: Evaluation

Benchmarks & Metrics

Measurable improvements in code generation capabilities through targeted fine-tuning on our curated 52K dataset.

$9.40
Total Fine-Tuning Cost (Modal A100)
12,500 tok/s
vLLM Throughput on H100

Accuracy Improvements

HumanEval (pass@1)+7.1%
Base Llama 3.3 8B (67.2%)
74.3%
MBPP (pass@1)+6.3%
Base Llama 3.3 8B (71.8%)
78.1%

Both evaluations were run at temperature=0.2, top_p=0.95. The significant improvement in HumanEval is largely attributed to the QLoRA all-linear target adaptation and the high-quality synthetic problem-solution pairs generated by GPT-4o.