🔁Export aus PyTorch – die Konvertierung
torch.nn.Module zum .mlpackage in elf Schritten – mit den echten Zwischenergebnissen der Konvertierung des Lehrmodells (coremltools 9.0, torch 2.7.0). 📎 PyTorch Conversion Workflow📎 eigene Konvertierung🐞Konvertierungs-Debugger
1. PyTorch-Modell
model = DigitMLP() # 64 → 32 → ReLU → 10 → Softmax
# … trainiert auf load_digits (Testgenauigkeit 96,7 %)Ausgangspunkt ist ein ganz normales torch.nn.Module. Für diese App: ein winziges MLP mit 2 410 Parametern, trainiert auf den 8×8-Ziffern aus scikit-learn.
torch.manual_seed(0), Adam lr=0.01, 300 Epochen Full-Batch, Cross-Entropy. Datensatz: scikit-learn load_digits (UCI Optical Recognition of Handwritten Digits, 8×8, Werte 0–16).
🔀PyTorch-Graph links, MIL rechts
Zeile anklicken für die Erklärung. Bei FLOAT32 ist die MIL-Spalte exakt die Op-Folge des konvertierten Programms; bei FLOAT16 kommen die cast-Ops dazu (im echten Programm stehen die Konstanten teils an anderer Stelle – siehe Debugger Schritt 7). Alle MIL-Ops: 📎 API-Referenz: MIL Ops📎 eigene Konvertierung
torch.jit.script wird nur eingeschränkt unterstützt. 📎 PyTorch Conversion Workflow🧾Der komplette Code
Kompletter Weg: PyTorch → .mlpackage (so für diese App ausgeführt)
✔ echte coremltools-Ausgabeimport numpy as np
import torch
import coremltools as ct # 9.0
class DigitMLP(torch.nn.Module):
def __init__(self):
super().__init__()
self.fc1 = torch.nn.Linear(64, 32)
self.fc2 = torch.nn.Linear(32, 10)
def forward(self, x):
return torch.softmax(self.fc2(torch.relu(self.fc1(x))), dim=-1)
model = DigitMLP()
# … Training …
model.eval() # 1. Dropout/BatchNorm in den Auswertungsmodus
example = torch.rand(1, 64) # 2. Beispiel-Eingabe mit der echten Form
traced = torch.jit.trace(model, example) # 3. TorchScript per Tracing
mlmodel = ct.convert( # 4. Konvertieren
traced,
inputs=[ct.TensorType(name="pixels", shape=(1, 64), dtype=np.float32)],
classifier_config=ct.ClassifierConfig([str(i) for i in range(10)]),
convert_to="mlprogram", # Standard seit coremltools 7.0
minimum_deployment_target=ct.target.iOS17,
compute_precision=ct.precision.FLOAT16, # Standard bei mlprogram
)
# 5. Metadaten – erscheinen in Xcode
mlmodel.author = "Visuelle Erklärungen (Beispielmodell)"
mlmodel.short_description = "Erkennt Ziffern 0–9 aus 8×8 Pixeln."
mlmodel.version = "1.0"
mlmodel.license = "Nur Lehrzwecke"
mlmodel.input_description["pixels"] = "64 Pixelwerte 0…1"
mlmodel.output_description["classLabel"] = "Wahrscheinlichste Ziffer"
mlmodel.save("DigitMLP.mlpackage") # 6. Speichern (ML Program → nur .mlpackage)Variante mit torch.export (ExportedProgram)
✔ echte coremltools-Ausgabeexample_inputs = (torch.rand(1, 64),)
ep = torch.export.export(model.eval(), example_inputs)
ep = ep.run_decompositions({}) # TRAINING → ATEN-Dialekt (sonst NotImplementedError)
mlmodel = ct.convert(
ep, # kein inputs= nötig: Formen stehen im ExportedProgram
classifier_config=ct.ClassifierConfig([str(i) for i in range(10)]),
minimum_deployment_target=ct.target.iOS17,
)inputs:ct.TensorType(→ MLMultiArray) oderct.ImageType(→ CVPixelBuffer/Bild)outputs: Namen/Typen der Ausgänge, auchImageType(seit coremltools 6)convert_to:"mlprogram"(Standard ab 7.0) oder"neuralnetwork"minimum_deployment_target:ct.target.iOS15 … iOS18, iOS26bzw. macOS-Pendantscompute_precision:ct.precision.FLOAT16(Standard) oderFLOAT32classifier_config:ct.ClassifierConfig(labels)compute_units: nur für das Laden in Python (predict), StandardALL
🖼️Bilder als Eingang: ImageType und Normalisierung
Werte laut coremltools-Guide „Image Input and Output“ für alle vortrainierten torchvision-Modelle.
| Kanal | mean | std |
|---|---|---|
| R | ||
| G | ||
| B |
Core ML: y = x · scale + bias
⇒ scale = 1/(255·std) · bias = −mean/std
| Kanal | exakter scale | bias |
|---|---|---|
| R | 0,01712475 | -2,11790393 |
| G | 0,017507 | -2,03571429 |
| B | 0,01742919 | -1,80444444 |
ct.ImageType(name="image", shape=(1, 3, H, W),
scale=1/(0.226*255.0),
bias=[-0.485/0.229, -0.456/0.224, -0.406/0.225],
color_layout=ct.colorlayout.RGB)Bildeingang mit Normalisierung (torchvision → ImageType)
✔ echte coremltools-Ausgabe# torchvision: y = (x/255 − mean) / std → Core ML: y = x · scale + bias
scale = 1 / (0.226 * 255.0) # ein globaler scale (Mittel der drei std)
bias = [-0.485 / 0.229, -0.456 / 0.224, -0.406 / 0.225]
mlmodel = ct.convert(
traced_cnn,
inputs=[ct.ImageType(name="image", shape=(1, 3, 224, 224),
scale=scale, bias=bias,
color_layout=ct.colorlayout.RGB)],
minimum_deployment_target=ct.target.iOS17,
)Die ersten Ops sind const → mul → const → add: erst mal scale, dann plus bias – als Float32-Hex-Konstanten.
program(1.0)
[buildInfo = dict<tensor<string, []>, tensor<string, []>>({{"coremlc-component-MIL", "3600.16.1"}, {"coremlc-version", "3600.25.2"}})]
{
func main<ios17>(tensor<fp32, [1, 3, 32, 32]> image) {
tensor<fp32, []> image__scaled___y_0 = const()[name = tensor<string, []>("image__scaled___y_0"), val = tensor<fp32, []>(0x1.1c4bep-6)];
tensor<fp32, [1, 3, 32, 32]> image__scaled__ = mul(x = image, y = image__scaled___y_0)[name = tensor<string, []>("image__scaled__")];
tensor<fp32, [1, 3, 1, 1]> image__biased___y_0 = const()[name = tensor<string, []>("image__biased___y_0"), val = tensor<fp32, [1, 3, 1, 1]>([[[[-0x1.0f177ap+1]], [[-0x1.04924ap+1]], [[-0x1.cdf012p+0]]]])];
tensor<fp32, [1, 3, 32, 32]> image__biased__ = add(x = image__scaled__, y = image__biased___y_0)[name = tensor<string, []>("image__biased__")];{"name":"image","type":{"imageType":{"width":"32","height":"32","colorSpace":"RGB"}}}📐Flexible Formen
Flexible Formen: RangeDim und EnumeratedShapes
✔ echte coremltools-Ausgabe# Batchgröße 1 … 64 (unbeschränkt ist bei mlprogram nicht erlaubt)
batch = ct.RangeDim(lower_bound=1, upper_bound=64, default=1)
flex = ct.convert(traced, inputs=[ct.TensorType(name="pixels", shape=(batch, 64))],
convert_to="mlprogram", minimum_deployment_target=ct.target.iOS17)
# Nur bestimmte Formen – laut Guide die schnellere Wahl (bis 128 Formen)
shapes = ct.EnumeratedShapes(shapes=[[1, 64], [8, 64], [32, 64]], default=[1, 64])
enum = ct.convert(traced, inputs=[ct.TensorType(name="pixels", shape=shapes)],
convert_to="mlprogram", minimum_deployment_target=ct.target.iOS17)
# Beobachtet: ohne dtype= wird der Eingang hier FLOAT16 (siehe Kapitel „Tipps & Fehler“){
"name": "pixels",
"type": {
"multiArrayType": {
"shape": [
"1",
"64"
],
"dataType": "FLOAT16",
"shapeRange": {
"sizeRanges": [
{
"lowerBound": "1",
"upperBound": "64"
},
{
"lowerBound": "64",
"upperBound": "64"
}
]
}
}
}
}{
"name": "pixels",
"type": {
"multiArrayType": {
"shape": [
"1",
"64"
],
"dataType": "FLOAT16",
"enumeratedShapes": {
"shapes": [
{
"shape": [
"1",
"64"
]
},
{
"shape": [
"8",
"64"
]
},
{
"shape": [
"32",
"64"
]
}
]
}
}
}
}