FrothSeg: run the same list on real images and the order inverts
Published:
FrothSeg: run the same list on real images and the order inverts

A follow-up finding in FrothSeg, and the reason its benchmark page carries two rankings instead of one.
On synthetic froth with exact per-bubble ground truth, fifteen methods sit in a certain order: the learned LamellaStar ensemble leads at mean AP 0.5186, just ahead of Cellpose-SAM at 0.5099, with the classical watershed family below. Then I took the same fifteen methods, applied the synthetic-fitted post-processing unchanged, and ran them over 64 real photographs of dense touching instances (BBBC038, CC0). Nothing retrained, nothing recalibrated. The numbers measure transfer, not fit.
Methods trained only on my generator fall the most (mean AP change -0.243). Classical methods with no learned prior move up on average (+0.071). The one method carrying a large external prior, Cellpose-SAM, rises (+0.199), because the adjacent domain plays to its pretraining. LamellaStar drops 0.394.
The load-bearing caveats stay attached: BBBC038 is an adjacent domain (touching nuclei, not froth), it plays to Cellpose-SAM’s cell-microscopy pretraining and says nothing about froth accuracy, and there is still no real froth data, so this cannot clear the release gate. It measures one thing only, whether the ladder generalises off the generator, and the answer for the models that only saw the generator is no. Two classical methods that move against their tier are named rather than averaged away. That is why the page shows both rankings: a method that wins on your own synthetic data can be the one that travels worst. Part of the Faena hub. Live · source.
