Palm infection screening - evaluation report

report 0.1.0_20260909-185125 · model 0.1.0+efficientnet_b0.20260909185105 · generated 20260909-185125 · config fb02ca6c3c098c5b

Limitations - read before the numbers

1. Dataset and splits

splitimagesindependent groupsinfectedhealthygroups (infected)groups (healthy)
dev160.000133.00071.00089.00061.00072.000
test40.00022.00018.00022.0009.00013.000

Grouping strategy: metadata_column (column: tree_id, trust: strong, 155 groups)

Test-set construction: group_stratified_random_holdout. Test groups were drawn at random from the same collection. Unseen trees, yes; unseen farm/camera/season, no. Performance on a genuinely new site will typically be lower.

Development folds

foldn_imagesn_groupsn_infected
032.00026.00014.000
132.00027.00014.000
232.00027.00015.000
332.00026.00014.000
432.00027.00014.000

2. Duplicate and leakage audit

exact duplicate groupsnear-duplicate clusterscross-class near duplicatesconflicting labels on identical files
20.00020.0000.0000.000

Possible shortcuts

fieldsingle-class valuesmessage
camera_id2.0002 value(s) of 'camera_id' map to (almost) one class only. The model can learn the acquisition condition instead of the disease.
site_id2.0002 value(s) of 'site_id' map to (almost) one class only. The model can learn the acquisition condition instead of the disease.
capture_date3.0003 value(s) of 'capture_date' map to (almost) one class only. The model can learn the acquisition condition instead of the disease.
session_id8.0008 value(s) of 'session_id' map to (almost) one class only. The model can learn the acquisition condition instead of the disease.
image_resolution3.000Some exact resolutions occur in one class only, which usually means the two classes came from different devices or sources.
Manual review checklist (a human must answer these)
  • Is the diseased tissue actually visible in the frame, or is the palm a distant speck?
  • Do infected and healthy photos come from different farms, seasons, or times of day?
  • Are there watermarks, logos, date stamps, or borders on one class only?
  • Do multiple palms appear in one frame, so the label is ambiguous?
  • Were infected images downloaded from the web while healthy ones were shot on a phone?
  • Do backgrounds differ systematically (sand vs irrigation, sky vs buildings)?
  • Are labels assigned per-tree by an agronomist, or guessed from appearance?

Label inspection montages

Contact sheets were written to reports/audit/20260909-172221: montage_healthy.png, montage_infected.png. Open them and check the labels by eye: no metric in this report can detect a systematically mislabelled class.

3. Development cross-validation DEVELOPMENT - NOT PERFORMANCE

OOF ROC-AUC
0.976
OOF PR-AUC
0.975
95% CI 0.943 - 0.997
group_bootstrap_percentile
dev images
160.000
dev groups
133.000

Per-fold results

foldn_valn_val_groupsbest_epochroc_aucpr_aucsensitivity@0.5specificity@0.5deviceseconds
0.00032.00026.0009.0000.9920.9900.8570.944mps213.700
1.00032.00027.0006.0001.0001.0000.8571.000mps174.300
2.00032.00027.0000.0001.0001.0000.6001.000mps162.600
3.00032.00026.0001.0001.0001.0000.9291.000mps161.500
4.00032.00027.0006.0000.9960.9950.9290.944mps174.500

Fold-to-fold spread

metricmeanstdminmax
roc_auc0.9980.0040.9921.000
pr_auc0.9970.0040.9901.000
balanced_accuracy0.9060.0630.8000.964
sensitivity_infected0.8340.1360.6000.929
specificity_healthy0.9780.0300.9441.000
development ROCdevelopment PRthreshold sweepscore distribution

Learning curves

learning curves

5. Calibration and decision threshold

calibration methodfitted onn calibration pointsthresholdpolicytarget recallselected onreview band
NoneNoneNone0.588max_specificity_at_recall0.900dev_out_of_fold0.100

6. Locked test set THE ONLY EVALUATION

Evaluated once on 40 images from 22 independent groups at threshold 0.588 (fixed before this evaluation). Class counts: {'healthy': 22, 'infected': 18}.

sensitivity (infected)
100.0%
95% CI 82.4% - 100.0%
wilson
false negative rate
0.0%
specificity (healthy)
100.0%
95% CI 85.1% - 100.0%
wilson
precision (infected)
100.0%
95% CI 82.4% - 100.0%
wilson
NPV (healthy)
100.0%
95% CI 85.1% - 100.0%
wilson
balanced accuracy
100.0%
95% CI 100.0% - 100.0%
group_bootstrap_percentile
macro F1
100.0%
95% CI 100.0% - 100.0%
group_bootstrap_percentile
MCC
1.000
95% CI 1.000 - 1.000
group_bootstrap_percentile
ROC-AUC
1.000
95% CI 1.000 - 1.000
group_bootstrap_percentile
PR-AUC
1.000
95% CI 1.000 - 1.000
group_bootstrap_percentile
log loss
0.017
Brier score
0.003
95% CI 0.000 - 0.007
group_bootstrap_percentile
ECE
0.016

Confusion matrix (raw counts)

predicted healthypredicted infected
actual healthy22.0000.000
actual infected0.00018.000

0 infected palm(s) were missed and 0 healthy palm(s) were flagged. 0 prediction(s) fell inside the review band and would be routed to a human.

confusion matrixtest ROCtest PRcalibrationscore distribution

7. What this means at real field prevalence

Sensitivity and specificity do not change with prevalence; precision and NPV do. This dataset is roughly balanced, a real grove is not.

field prevalenceprecision (PPV)NPVfalse alarms per 100 healthymissed per 100 infected
1.0%1.0001.0000.0000.000
5.0%1.0001.0000.0000.000
10.0%1.0001.0000.0000.000
25.0%1.0001.0000.0000.000
50.0%1.0001.0000.0000.000

9. Reproducibility

seedsplit manifestconfig fingerprintmodel versionweights sha256deviceplatformgit commit
1337.00020260909-181039_seed1337_fb02ca6c3c098c5bfb02ca6c3c098c5b0.1.0+efficientnet_b0.2026090918510596a80aa9ddd7252ampsmacOS-27.0-arm64-arm-64bit-Mach-Onot a git checkout
Package versions
packageversion
python3.13.0
torch2.14.0
torchvision0.29.0
numpy2.5.3
sklearn1.9.0
PIL12.3.0
scipy1.18.1
pandas3.0.5
Full configuration
{
  "project_name": "palm-screen",
  "model_version": "0.1.0",
  "data": {
    "raw_dir": "data/raw",
    "metadata_csv": "data/metadata.csv",
    "classes": [
      "healthy",
      "infected"
    ],
    "positive_class": "infected",
    "group_key_priority": [
      "tree_id",
      "session_id",
      "site_id"
    ],
    "phash_pseudo_groups": true,
    "phash_group_distance": 6,
    "allow_image_level_split": true,
    "holdout_column": null,
    "holdout_values": []
  },
  "split": {
    "test_fraction": 0.2,
    "n_folds": 5,
    "seed": 1337,
    "manifest_dir": "artifacts/splits"
  },
  "train": {
    "arch": "efficientnet_b0",
    "image_size": 224,
    "batch_size": 16,
    "num_workers": 2,
    "head_epochs": 6,
    "finetune_epochs": 25,
    "lr_head": 0.001,
    "lr_backbone": 0.0001,
    "weight_decay": 0.0001,
    "label_smoothing": 0.0,
    "unfreeze_last_n_blocks": 3,
    "freeze_batchnorm": true,
    "early_stopping_metric": "val_pr_auc",
    "early_stopping_patience": 6,
    "early_stopping_min_delta": 0.0001,
    "grad_clip_norm": 1.0,
    "use_class_weights": false,
    "amp": false,
    "seed": 1337,
    "checkpoint_dir": "artifacts/checkpoints",
    "augmentation": {
      "random_resized_crop": true,
      "crop_scale_min": 0.8,
      "crop_scale_max": 1.0,
      "rotation_degrees": 12,
      "horizontal_flip": true,
      "vertical_flip": false,
      "brightness": 0.15,
      "contrast": 0.15,
      "saturation": 0.1,
      "hue": 0.02,
      "mixup": false,
      "cutmix": false
    }
  },
  "threshold": {
    "policy": "max_specificity_at_recall",
    "target_recall": 0.9,
    "fixed_threshold": 0.5,
    "review_band": 0.1
  },
  "calibration": {
    "method": "platt"
  },
  "report": {
    "output_dir": "reports",
    "bootstrap_iterations": 2000,
    "ece_bins": 10,
    "prevalence_grid": [
      0.01,
      0.05,
      0.1,
      0.25,
      0.5
    ]
  }
}
Split manifest (counts and leakage audit)
{
  "manifest_version": "20260909-181039_seed1337_fb02ca6c3c098c5b",
  "created_at": "20260909-181039",
  "seed": 1337,
  "config_fingerprint": "fb02ca6c3c098c5b",
  "grouping": {
    "strategy": "metadata_column",
    "column": "tree_id",
    "n_groups": 155,
    "trust": "strong",
    "warnings": []
  },
  "split_info": {
    "strategy": "group_stratified_random_holdout",
    "stats": {
      "target_n": 40,
      "actual_n": 40,
      "n_infected": 18,
      "n_healthy": 22,
      "forced_groups_for_class_balance": [],
      "actual_pos": 18,
      "actual_pos_rate": 0.45,
      "overall_pos_rate": 0.445
    },
    "note": "Test groups were drawn at random from the same collection. Unseen trees, yes; unseen farm/camera/season, no. Performance on a genuinely new site will typically be lower.",
    "n_dev": 160,
    "n_dev_infected": 71,
    "n_dev_groups": 133,
    "n_test": 40,
    "n_test_infected": 18,
    "n_test_groups": 22
  },
  "leakage_audit": {
    "clean": true,
    "group_overlap": [],
    "sha256_overlap": [],
    "near_duplicate_pairs_across_splits": [],
    "fold_group_overlap": []
  },
  "counts": {
    "total_images": 200,
    "dev": {
      "n_images": 160,
      "n_groups": 133,
      "by_class": {
        "healthy": 89,
        "infected": 71
      },
      "groups_by_class": {
        "healthy": 72,
        "infected": 61
      },
      "folds": {
        "0": {
          "n_images": 32,
          "n_groups": 26,
          "n_infected": 14
        },
        "1": {
          "n_images": 32,
          "n_groups": 27,
          "n_infected": 14
        },
        "2": {
          "n_images": 32,
          "n_groups": 27,
          "n_infected": 15
        },
        "3": {
          "n_images": 32,
          "n_groups": 26,
          "n_infected": 14
        },
        "4": {
          "n_images": 32,
          "n_groups": 27,
          "n_infected": 14
        }
      }
    },
    "test": {
      "n_images": 40,
      "n_groups": 22,
      "by_class": {
        "healthy": 22,
        "infected": 18
      },
      "groups_by_class": {
        "healthy": 13,
        "infected": 9
      }
    }
  },
  "csv": "split_20260909-181039_seed1337_fb02ca6c3c098c5b.csv",
  "provenance": {
    "platform": "macOS-27.0-arm64-arm-64bit-Mach-O",
    "machine": "arm64",
    "processor": "arm",
    "git_commit": null,
    "packages": {
      "python": "3.13.0",
      "torch": "2.14.0",
      "torchvision": "0.29.0",
      "numpy": "2.5.3",
      "sklearn": "1.9.0",
      "PIL": "12.3.0",
      "scipy": "1.18.1",
      "pandas": "3.0.5"
    },
    "device": {
      "device": "mps",
      "reason": "Apple Silicon GPU (MPS) available",
      "mps_available": true,
      "mps_built": true
    }
  }
}

10. Grad-CAM