add tvm-ffi metal - #506
Open
ved1beta wants to merge 1 commit into
Open
Conversation
Member
|
We already have a branch for tvm-ffi and there is also support for kernel-builder in it. However Metal support will require upstream tvm-ffi changes, because there is currently no way for tvm-ffi to pass For this reason we are postponing tvm-ffi Metal support until these issue have been fixed upstream. |
Author
|
ahh thanks for the response , will try to put a pr upstream 🫡 |
This was referenced May 1, 2026
pytorchmergebot
pushed a commit
to pytorch/pytorch
that referenced
this pull request
Aug 11, 2026
## Summary `toDLPackNonOwning` writes `src.data_ptr()` into `DLTensor.data` and sets `byte_offset = 0`. For MPS tensors this is wrong: PyTorch's MPS allocator encodes `id<MTLBuffer>` directly into `c10::DataPtr.data` in `aten/src/ATen/mps/MPSAllocator.mm:768-770`, so for any sliced/viewed MPS tensor `data_ptr()` returns `id<MTLBuffer> + storage_offset*elemsize` — pointer arithmetic on an Objective-C object pointer, which is not a valid buffer handle. The owning path already special-cases this correctly. The non-owning path was missing the same logic. ##fix mirror the owning path's MPS handling: write the storage base into DLTensor.data and the element offset into byte_offset ##Test Adds test_dlpack_exchange_api_mps_sliced in test/test_dlpack.py, which exercises the C exchange API on a sliced MPS tensor and asserts related upstreams apache/tvm-ffi#578 huggingface/kernels#506 thanks Pull Request resolved: #182924 Approved by: https://github.com/Skylion007 Co-authored-by: Aaron Gokaslan <aaronGokaslan@gmail.com>
ankitbhrdwj
pushed a commit
to ankitbhrdwj/pytorch
that referenced
this pull request
Aug 12, 2026
…2924) ## Summary `toDLPackNonOwning` writes `src.data_ptr()` into `DLTensor.data` and sets `byte_offset = 0`. For MPS tensors this is wrong: PyTorch's MPS allocator encodes `id<MTLBuffer>` directly into `c10::DataPtr.data` in `aten/src/ATen/mps/MPSAllocator.mm:768-770`, so for any sliced/viewed MPS tensor `data_ptr()` returns `id<MTLBuffer> + storage_offset*elemsize` — pointer arithmetic on an Objective-C object pointer, which is not a valid buffer handle. The owning path already special-cases this correctly. The non-owning path was missing the same logic. ##fix mirror the owning path's MPS handling: write the storage base into DLTensor.data and the element offset into byte_offset ##Test Adds test_dlpack_exchange_api_mps_sliced in test/test_dlpack.py, which exercises the C exchange API on a sliced MPS tensor and asserts related upstreams apache/tvm-ffi#578 huggingface/kernels#506 thanks Pull Request resolved: pytorch#182924 Approved by: https://github.com/Skylion007 Co-authored-by: Aaron Gokaslan <aaronGokaslan@gmail.com>
gavinwang269
pushed a commit
to gavinwang269/pytorch
that referenced
this pull request
Aug 13, 2026
…2924) ## Summary `toDLPackNonOwning` writes `src.data_ptr()` into `DLTensor.data` and sets `byte_offset = 0`. For MPS tensors this is wrong: PyTorch's MPS allocator encodes `id<MTLBuffer>` directly into `c10::DataPtr.data` in `aten/src/ATen/mps/MPSAllocator.mm:768-770`, so for any sliced/viewed MPS tensor `data_ptr()` returns `id<MTLBuffer> + storage_offset*elemsize` — pointer arithmetic on an Objective-C object pointer, which is not a valid buffer handle. The owning path already special-cases this correctly. The non-owning path was missing the same logic. ##fix mirror the owning path's MPS handling: write the storage base into DLTensor.data and the element offset into byte_offset ##Test Adds test_dlpack_exchange_api_mps_sliced in test/test_dlpack.py, which exercises the C exchange API on a sliced MPS tensor and asserts related upstreams apache/tvm-ffi#578 huggingface/kernels#506 thanks Pull Request resolved: pytorch#182924 Approved by: https://github.com/Skylion007 Co-authored-by: Aaron Gokaslan <aaronGokaslan@gmail.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
fixes #352
Summary
tvm-ffi-metal,tvm-ffi-cuda,tvm-ffi-cpu,tvm-ffi-rocm,tvm-ffi-xpu,tvm-ffi-universal)TvmFfiNoarchmirroringTorchNoarch;parse_variantandBUILD_VARIANT_REGEX_sort_variantsso within noarch, Torch is preferred over tvm-ffitests
Torch beats tvm-ffi when both are present
noarch fallback when no arch variant matches
arch variant beats noarch on the actual hardware