A 1.3 billion parameter model matching GPT-4 on text-to-SQL benchmarks. A fine-tuned 7B model beating ChatGPT on tool-calling by 3x. A healthcare NLP system hitting 96% accuracy where GPT-4o manages 79%.

These aren’t cherry-picked outliers.

Research across multiple studies found that fine-tuned...