What changed, and when.
Candela's own history and developments, separate from Lantern's product changelog.
candela-2.2
The serving model advanced again, and for the first time the new version is better than its predecessor on every benchmark we run: art recall rose to 84% from 83% at the same 1% false-alarm budget, and the photography score set a program record at 90% of the leading general-purpose model.
Two training changes made the trade unnecessary: the crop-matching supervision no longer taxes photographs, and photographs are actively held near the base model's understanding while the art specialization deepens. Scoring thresholds and bands are unchanged.
candela-2.1
The serving model advanced to a newer generation of the same 22M network, selected against the same benchmarks before the swap. Scoring thresholds and bands are unchanged.
On our art benchmark it catches 83% of altered copies at the same 1% false-alarm budget, up from 76%, with the gains concentrated in crops and rotations and no family regressing. Its photography-benchmark score remains more than double the retired ensemble's.
candela-2-pilot
The first version that is Lantern's own model: a single 22M-parameter network fine-tuned for art, replacing the three-model ensemble. One 512-number fingerprint per artwork.
On our art benchmark it catches 76% of altered copies at a 1% false-alarm budget, ahead of every open model we have measured and within one point of the far larger ensemble it replaced, while more than doubling the ensemble's score on the photography benchmark.
candela-1.3
A newly released open copy-detection model joined the ensemble as a third scored signal, with the fusion weights re-derived on our evaluation instrument. Recall on the art benchmark moved from 74% to 77% at the same false-alarm budget.
candela-1.2
Every candidate is now scored on the same fixed set of signals instead of whichever signals happened to retrieve it. The variable basis had been flattering unrelated works and diluting true matches; fixing it cut the false-flag rate at benchmark scale from 96% to 5% of unrelated works.
candela-1.1
A decision-rule change rather than a new model: a weak match now has to survive on a signal other than the vision-language model, which is the one most prone to confusing similar style with the same work.
Cuts false positives substantially against the previous rule at effectively no cost in recall.
candela-1
The first production system: three separate models fused with weights tuned on our evaluation set. Superseded by candela-1.1.