Compress large language models using knowledge distillation from teacher to student models. Use when deploying smaller …
Compress large language models using knowledge distillation from teacher to student models. Use when deploying smaller models with retained performance, transferring GPT-4 capabilities to open-source models, or reducing inference costs. Covers temperature scaling, soft targets, reverse KLD, logit distillation, and MiniLLM training strategies.
davila7
cli
free
Others in the same category, ranked by how often they are opened.