Model card

What PictureCloud knows—and where it falls short

PictureCloud offers four locally runnable classification modes. This page records their scope and contracts, with measured performance shown only where it has been independently preserved.

At a glance

46.25%TFLite test accuracy
0.4802test macro-F1
2.0 MBmodel file
11output labels

Those results make the artifact useful for engineering, comparison, and experimentation—not for authoritative identification. PictureCloud deliberately calls outputs “model scores” and shows alternatives rather than converting a score into certainty.

Four selectable detail levels

ModeVocabularyInterpretation
Experimental11 labelsCompact baseline: ten genera plus Contrail
Broad12 labelsMain cloud types plus Clear sky
Detailed15 labelsCloud species and structural forms
Fine detail25 scoresIndependent features and varieties; several may apply

The Broad, Detailed, and Fine detail models use a 256 × 256 input and standardized RGB values. Their scores are available for visual comparison, but PictureCloud does not publish an independent accuracy result for them. The 46.25% figure above applies only to the Experimental model.

Dataset and split

The training notebook audited 2,543 readable RGB images in CCSN v2. It excluded exact duplicates with conflicting labels and retained one copy from each same-class duplicate pair, leaving 2,523 images. A deterministic stratified split produced 2,018 training, 252 validation, and 253 held-out test images.

The model vocabulary follows the dataset folder order: Altocumulus, Altostratus, Cumulonimbus, Cirrocumulus, Cirrus, Cirrostratus, Contrail, Cumulus, Nimbostratus, Stratocumulus, and Stratus. Contrail is a dataset class, not one of the ten WMO cloud genera.

Provenance and release status

The bundled file is the FP16 export of the project’s CCSN-trained MobileNetV3Small checkpoint. The dataset is the Cirrus Cumulus Stratus Nimbus Database, DOI 10.7910/DVN/CADDPD, associated with Zhang, Liu, Zhang, and Song’s 2018 CloudNet paper.

Provenance record needs completionThe public products currently contain this model, but the local release record still lacks preserved evidence of the DOI-specific dataset terms and an explicit trained-weight redistribution decision. The operator should verify and archive those terms. This page records the gap and does not itself grant redistribution rights.

Evaluation

The selected Keras checkpoint measured 47.04% accuracy and 0.4873 macro-F1 on the held-out split. Its FP16 TensorFlow Lite export measured 46.25% accuracy and 0.4802 macro-F1, with 99.21% predicted-class agreement between Keras and TFLite.

Performance varies substantially by label. Contrails and Cumulonimbus were among the stronger test classes; Stratus, Nimbostratus, and Cirrostratus were important weaknesses. A single aggregate percentage must not hide those uneven errors.

Browser and mobile contract

BoundaryContract
Input[1, 224, 224, 3] float32 RGB
Pixel valuesRaw 0–255; MobileNetV3 normalization is embedded
Website cropOrientation-corrected centered square, then direct resize
Output[1, 11] float32 softmax scores
CompressionFP16 weights with float32 input and output
RuntimeBuilt-in TFLite operators through LiteRT.js WASM

Known limitations

  • No Clear or unknown class: every input receives one of eleven labels, including empty sky, indoor photos, landscapes, or corrupted visual context.
  • No height or motion: a photograph removes two of the most useful identification clues.
  • Center crop: a wide pattern or halo can fall outside the model view.
  • Mixed skies: the model produces one ranked distribution even when several genera coexist.
  • Dataset shift: cameras, regions, seasons, exposure, and atmospheric conditions outside CCSN can reduce performance.
  • Not calibrated for decisions: a softmax value is not a guaranteed probability of correctness.

What the tool does with uncertainty

The result view exposes the top three and all eleven scores. It highlights low leading scores or narrow margins, reports basic resolution and exposure diagnostics, repeats the missing-Clear limitation, and links to human-readable visual cues. None of those safeguards turns the baseline into a forecasting or warning system.

Production recommendationKeep the app and website model-swappable. A future candidate should be evaluated on the same held-out split plus broader external data, include an out-of-domain strategy, and pass Keras↔TFLite↔browser parity tests before release.

Taxonomy references

Search PictureCloud