ashxhart commited on
Commit
9872546
·
verified ·
1 Parent(s): 9938ce0

Rebrand model card to TensorFold

Browse files
Files changed (1) hide show
  1. README.md +15 -10
README.md CHANGED
@@ -26,6 +26,11 @@ tags:
26
  - apple-silicon
27
  ---
28
 
 
 
 
 
 
29
 
30
  <div align="center">
31
  <a href="https://huggingface.co/zai-org/GLM-5.3-Flash">
@@ -38,7 +43,7 @@ tags:
38
  <img src="https://img.shields.io/badge/Z.ai-GLM--5.3--Flash-111827?style=for-the-badge" alt="Z.ai GLM-5.3-Flash">
39
  <img src="https://img.shields.io/badge/Apple_Silicon-MLX-000000?style=for-the-badge&logo=apple&logoColor=white" alt="Apple silicon MLX">
40
  <img src="https://img.shields.io/badge/Native_MTP-Included-22C55E?style=for-the-badge" alt="Native MTP included">
41
- <img src="https://img.shields.io/badge/Vontra-oQ-6E56CF?style=for-the-badge&logo=huggingface&logoColor=white" alt="Vontra oQ">
42
  </p>
43
 
44
 
@@ -65,7 +70,7 @@ tags:
65
  | Item | Value |
66
  | --- | --- |
67
  | Base model | [`zai-org/GLM-5.3-Flash`](https://huggingface.co/zai-org/GLM-5.3-Flash) |
68
- | Repository | `Vontra/GLM-5.3-Flash-MLX-oQ4-MTP` |
69
  | Format | MLX safetensors |
70
  | Quantisation | oQ4 mixed precision: 4-bit affine base with 554 sensitivity-selected 5/6/8-bit overrides |
71
  | Base group size | 64 |
@@ -97,7 +102,7 @@ The upstream tokenizer, chat template, multimodal processor, generation configur
97
  | Vision encoder and projector | source-compatible precision |
98
  | Native MTP prediction layer | 4-bit affine base with 12 native-MTP overrides at 5/6/8-bit |
99
  | Other non-quantisable tensors | Preserved at source-compatible precision |
100
- | Converter | Vontra streamed oQ converter using MLX 0.32.0 |
101
 
102
 
103
  Sensitivity was measured with the built-in multilingual code calibration set, 128 samples at 256 tokens. The allocation rule was byte-budgeted layer-sensitivity ranking under the oQ4 target and hard cap. These details are part of the release recipe and should be used when comparing oQ variants.
@@ -128,7 +133,7 @@ GLM-5.3-Flash uses the new `glm5_next` multimodal architecture, hybrid linear an
128
 
129
 
130
  ```bash
131
- hf download Vontra/GLM-5.3-Flash-MLX-oQ4-MTP \
132
  --local-dir GLM-5.3-Flash-MLX-oQ4-MTP
133
  ```
134
 
@@ -236,7 +241,7 @@ This is a community quantisation and is not an official Z.ai release.
236
  The upstream model is released under the **MIT License**. The required licence text is included in this repository.
237
 
238
 
239
- Model design, training, upstream evaluations, and documentation belong to Z.ai and the GLM-5 contributors. The oQ conversion, Apple-silicon validation, native-MTP integration work, and packaging are provided by [Vontra](https://huggingface.co/Vontra).
240
 
241
 
242
  If you use this model in research, cite the upstream report:
@@ -256,10 +261,10 @@ If you use this model in research, cite the upstream report:
256
 
257
 
258
 
259
- <!-- vontra-chooser-start -->
260
  ## Choose for your Mac
261
 
262
- [64GB Macs](https://huggingface.co/collections/Vontra/mlx-models-for-64gb-macs-6a9fefda17932216ec9ab457) · [128GB Macs](https://huggingface.co/collections/Vontra/mlx-models-for-128gb-macs-6a9ff0abd31bc9abbe7922d7) · [256GB Macs](https://huggingface.co/collections/Vontra/mlx-models-for-256gb-macs-6a9ff0ef9fed7c5bdca15e9b)
263
 
264
  No measured memory tier is assigned here. The collections use published M3 Studio peaks with at least 25% nominal headroom; fit on other Macs is an estimate, and full context is not guaranteed. Start with short context and one request.
265
 
@@ -270,7 +275,7 @@ oMLX version mentioned in the existing card: **0.6.3rc3**; consult its compatibi
270
  ### Quick start and demo prompt
271
 
272
  ```bash
273
- hf download Vontra/GLM-5.3-Flash-MLX-oQ4-MTP --local-dir ./models/GLM-5.3-Flash-MLX-oQ4-MTP
274
  ```
275
 
276
  Add the downloaded folder to oMLX model directories, refresh the list, and follow this card's architecture and MTP compatibility requirements before loading.
@@ -283,5 +288,5 @@ Explain why the sky looks blue in three short sentences.
283
 
284
  This is a demo prompt to try, not a recorded successful run; a captured demonstration for this documentation update is not yet available.
285
 
286
- [Follow Vontra for new Apple Silicon releases and fixes.](https://huggingface.co/Vontra)
287
- <!-- vontra-chooser-end -->
 
26
  - apple-silicon
27
  ---
28
 
29
+ <p align="center">
30
+ <a href="https://tensorfold.dev">
31
+ <img src="https://huggingface.co/spaces/TensorFold/README/resolve/main/tensorfold-logo.png" alt="TensorFold" width="160">
32
+ </a>
33
+ </p>
34
 
35
  <div align="center">
36
  <a href="https://huggingface.co/zai-org/GLM-5.3-Flash">
 
43
  <img src="https://img.shields.io/badge/Z.ai-GLM--5.3--Flash-111827?style=for-the-badge" alt="Z.ai GLM-5.3-Flash">
44
  <img src="https://img.shields.io/badge/Apple_Silicon-MLX-000000?style=for-the-badge&logo=apple&logoColor=white" alt="Apple silicon MLX">
45
  <img src="https://img.shields.io/badge/Native_MTP-Included-22C55E?style=for-the-badge" alt="Native MTP included">
46
+ <img src="https://img.shields.io/badge/TensorFold-oQ-6E56CF?style=for-the-badge&logo=huggingface&logoColor=white" alt="TensorFold oQ">
47
  </p>
48
 
49
 
 
70
  | Item | Value |
71
  | --- | --- |
72
  | Base model | [`zai-org/GLM-5.3-Flash`](https://huggingface.co/zai-org/GLM-5.3-Flash) |
73
+ | Repository | `TensorFold/GLM-5.3-Flash-MLX-oQ4-MTP` |
74
  | Format | MLX safetensors |
75
  | Quantisation | oQ4 mixed precision: 4-bit affine base with 554 sensitivity-selected 5/6/8-bit overrides |
76
  | Base group size | 64 |
 
102
  | Vision encoder and projector | source-compatible precision |
103
  | Native MTP prediction layer | 4-bit affine base with 12 native-MTP overrides at 5/6/8-bit |
104
  | Other non-quantisable tensors | Preserved at source-compatible precision |
105
+ | Converter | TensorFold streamed oQ converter using MLX 0.32.0 |
106
 
107
 
108
  Sensitivity was measured with the built-in multilingual code calibration set, 128 samples at 256 tokens. The allocation rule was byte-budgeted layer-sensitivity ranking under the oQ4 target and hard cap. These details are part of the release recipe and should be used when comparing oQ variants.
 
133
 
134
 
135
  ```bash
136
+ hf download TensorFold/GLM-5.3-Flash-MLX-oQ4-MTP \
137
  --local-dir GLM-5.3-Flash-MLX-oQ4-MTP
138
  ```
139
 
 
241
  The upstream model is released under the **MIT License**. The required licence text is included in this repository.
242
 
243
 
244
+ Model design, training, upstream evaluations, and documentation belong to Z.ai and the GLM-5 contributors. The oQ conversion, Apple-silicon validation, native-MTP integration work, and packaging are provided by [TensorFold](https://huggingface.co/TensorFold).
245
 
246
 
247
  If you use this model in research, cite the upstream report:
 
261
 
262
 
263
 
264
+ <!-- TensorFold-chooser-start -->
265
  ## Choose for your Mac
266
 
267
+ [64GB Macs](https://huggingface.co/collections/TensorFold/mlx-models-for-64gb-macs-6a9fefda17932216ec9ab457) · [128GB Macs](https://huggingface.co/collections/TensorFold/mlx-models-for-128gb-macs-6a9ff0abd31bc9abbe7922d7) · [256GB Macs](https://huggingface.co/collections/TensorFold/mlx-models-for-256gb-macs-6a9ff0ef9fed7c5bdca15e9b)
268
 
269
  No measured memory tier is assigned here. The collections use published M3 Studio peaks with at least 25% nominal headroom; fit on other Macs is an estimate, and full context is not guaranteed. Start with short context and one request.
270
 
 
275
  ### Quick start and demo prompt
276
 
277
  ```bash
278
+ hf download TensorFold/GLM-5.3-Flash-MLX-oQ4-MTP --local-dir ./models/GLM-5.3-Flash-MLX-oQ4-MTP
279
  ```
280
 
281
  Add the downloaded folder to oMLX model directories, refresh the list, and follow this card's architecture and MTP compatibility requirements before loading.
 
288
 
289
  This is a demo prompt to try, not a recorded successful run; a captured demonstration for this documentation update is not yet available.
290
 
291
+ [Follow TensorFold for new Apple Silicon releases and fixes.](https://huggingface.co/TensorFold)
292
+ <!-- TensorFold-chooser-end -->