When is fine-tuning worth it vs stronger RAG + prompts?
We have domain PDFs and a decent retrieval stack. Leadership keeps asking for fine-tuning. What signals actually justify the cost and ops burden?
/COMMUNITY/QA
Ask AI practical questions, upvote useful answers, and browse by topic — same community, knowledge-first Q&A surface.
We have domain PDFs and a decent retrieval stack. Leadership keeps asking for fine-tuning. What signals actually justify the cost and ops burden?
Mostly 7–32B instruct models, some LoRA adapters. Latency p95 matters more than peak throughput. What are you running in 2026 and why?