Weighing mainstream and alternative accounts…
Two lenses on the same evidence. Source weight and the primary source ratio show what each rests on.
Deeper threads worth pulling on next.
Investigated
Image: maxim-blog.ghost.io
Distillation-based copying trains a student model on a teacher model’s outputs—sometimes probability distributions or logits, sometimes sampled responses—to reproduce capabilities more cheaply or in a smaller system. It can be a legitimate compression technique, but API-based extraction may be described as copying when outputs are collected to build a substitute model; the label alone does not determine whether infringement occurred.
Two lenses on the same evidence. Source weight and the primary source ratio show what each rests on.
Lens adapted to this topic: Replication, provenance, and disputed boundaries
A broader and more critical usage treats distillation-based copying as capability extraction from an accessible model, especially through large-scale API queries. This perspective emphasizes that students can reproduce distinctive behaviors or teacher-derived sentences, and that the boundary between ordinary learning, model replication, and unauthorized appropriation depends on access practices, provenance, and legal rights.
Deeper threads worth pulling on next.