GPT-3 has 175 billion parameters.

Full fine-tuning updates all 175 billion with every gradient step. You need multiple A100 GPUs (each with 80GB memory) just to fit the model. Training for even a few epochs on a moderate dataset costs thousands of dollars. A startup cannot do this. A PhD student cannot do this.

Yet fine-tuned versions of large...