bump torchao pin
CLA Signed
Performance improvements to lowbit quantized linear metal kernels in torchao. See [AO PR#2167](https://github.com/pytorch/ao/pull/2167) for details.
The table below summarizes torchchat's decode speed (tokens/second) on Metal backend on M1 Max 64GB after this update
| # bits | Llama 3.2-1B | Llama 3.2-3B | Llama 3.1-8B |
| --- | --- | --- | --- |
| 1 | 179.96 | 87.01 | 51.10 |
| 2 | 186.91 | 98.13 | 54.37 |
| 3 | 170.62 | 85.69 | 48.05 |
| 4 | 175.15 | 89.54 | 50.51 |
| 5 | 147.10 | 70.19 | 38.58 |
| 6 | 140.51 | 63.62 | 35.48 |
| 7 | 131.27 | 64.19 | 32.69 |
合并状态:未合并 2 条评论