Jay Yun
Speculative draft checkpoints and compact embedding exports. Find the right model by family, runtime and stored precision.
Speculative Decoding Models
Six DSpark, DFlash and DFlash2 drafts for MiniCPM5-SFT and Ling-3.0-tiny.
Browse collection ↗EmbeddingGemma 2
Text (270M) and text + image (440M). Transformers and MLX, BF16 and FP32.
Browse collection ↗EmbeddingGemma 2
Eight published exports. Public download and inference checks are recorded in each model repository.
| Runtime | Stored precision | Text + code · 270M | Text + image · 440M |
|---|---|---|---|
| Transformers / SentenceTransformers | BF16 | Text BF16 ↗ | Text + image BF16 ↗ |
| Transformers / SentenceTransformers | FP32 | Text FP32 ↗ | Text + image FP32 ↗ |
| MLX | BF16 | Text BF16 ↗ | Text + image BF16 ↗ |
| MLX | FP32 | Text FP32 ↗ | Text + image FP32 ↗ |
Precision labels describe stored weights. Check each card for runtime requirements and validation scope; 270M / 440M are rounded deployment sizes.
Earlier BF16 exports / compatibility references
Existing integrations: text-270m · text-image-440m. Both store BF16 weights and appear after the eight precision-specific exports in the collection.
Independent exports derived from Google DeepMind’s EmbeddingGemma 2. See model cards for attribution.
Speculative decoding
| Target model | Draft families |
|---|---|
| openbmb/MiniCPM5-2B-SFT | DSpark · DFlash CE · DFlash DPACE · DFlash2 |
| inclusionAI/Ling-3.0-tiny | DSpark · DSpark-50K base checkpoint |
Pair each draft with the target named in its card. Runtime instructions, source terms and evaluation scope are documented there.
Browse all six draft checkpoints ↗