Skip to content

xFormers memory-efficient attention not supported on CPU  #310

Description

@Wasiq1123

I’m trying to run Depth-Anything-V2 with xFormers on my system (CPU only).
I get the following error:

NotImplementedError: No operator found for memory_efficient_attention_forward with inputs:
query : shape=(1, 1531, 6, 64) (torch.float32)
key : shape=(1, 1531, 6, 64) (torch.float32)
value : shape=(1, 1531, 6, 64) (torch.float32)
attn_bias : <class 'NoneType'>
p : 0.0
fa3F@2.8.3-133-gde1584b is not supported because:
device=cpu (supported: {'cuda'})
dtype=torch.float32 (supported: {torch.bfloat16, torch.float16})

It seems that memory-efficient attention in xFormers requires CUDA.

My environment:

  • PyTorch 2.5.7 / 2.8.3
  • xFormers installed from pip
  • Running on CPU only (no GPU available)
  • Python 3.10, Ubuntu 22.04

Question:
Is there a way to run Depth-Anything-V2 on CPU without a GPU, or do I have to disable memory-efficient attention? How can I fix this error on CPU?

Below is My code

#If this file give error of importing package then run it in this directory /testing_model/depth_models/src/Depth-Anything-V2

import cv2
import torch
import sys
sys.path.append('/home/wasiq/testing_model/depth_models/src/Depth-Anything-V2')
from depth_anything_v2.dpt import DepthAnythingV2

model_configs = {
'vits': {'encoder': 'vits', 'features': 64, 'out_channels': [48, 96, 192, 384]},
'vitb': {'encoder': 'vitb', 'features': 128, 'out_channels': [96, 192, 384, 768]},
'vitl': {'encoder': 'vitl', 'features': 256, 'out_channels': [256, 512, 1024, 1024]}
}

encoder = 'vitb' # or 'vits', 'vitb'
dataset = 'hypersim' # 'hypersim' for indoor model, 'vkitti' for outdoor model
max_depth = 20 # 20 for indoor model, 80 for outdoor model

model = DepthAnythingV2(**{**model_configs[encoder], 'max_depth': max_depth})

model = DepthAnythingV2(**{**model_configs[encoder]})
model.load_state_dict(torch.load(f'/home/wasiq/testing_model/depth_models/src/Depth-Anything-V2/metric_depth/checkpoints/depth_anything_v2_metric_{dataset}_{encoder}.pth', map_location='cpu'))
model.eval()

raw_img = cv2.imread('your/image/path')
depth = model.infer_image(raw_img) # HxW depth map in meters in numpy

Activity

  1. changed the title [-]xFormers memory-efficient attention not supported on CPU : NotImplementedError: memory_efficient_attention_forward not supported on CPU with float32 in xFormers[/-] [+]xFormers memory-efficient attention not supported on CPU [/+] on Dec 3, 2025
  2. heyoeyo commented on Dec 5, 2025

    @heyoeyo

    The model includes checks for whether xformers is installed or not. If it isn't installed, then it should fall back to a normal (non-xformers) implementation, so just doing pip uninstall xformers should fix things.

  3. Wasiq1123 commented on Dec 6, 2025

    @Wasiq1123
    Author

    Hi @heyoeyo — thanks for the help earlier. Uninstalling xFormers removed the import error, but now I'm seeing the model produce an all-zero depth map.

    ### Environment

    • OS: Ubuntu 22.04
    • Python: 3.10
    • Device: CPU only (no CUDA; torch.device -> 'cpu')
    • xFormers: uninstalled (or disabled via XFORMERS_DISABLE_MEMORY_EFFICIENT_ATTENTION=1)
    • Repo version: Depth-Anything-V2 (local copy)

    ### Encoder / checkpoint

    • Encoder used: vits (I also tested vitb, vitg variants)
    • Checkpoint: metric_depth/checkpoints/depth_anything_v2_metric_hypersim_vits.pth (loaded with map_location='cpu')

    ### Minimal repro (my code)

    import cv2
    import torch
    from depth_anything_v2.dpt import DepthAnythingV2
    
    DEVICE = 'cuda' if torch.cuda.is_available() else 'mps' if torch.backends.mps.is_available() else 'cpu'
    
    model_configs = {
        'vits': {'encoder': 'vits', 'features': 64, 'out_channels': [48, 96, 192, 384]},
        'vitb': {'encoder': 'vitb', 'features': 128, 'out_channels': [96, 192, 384, 768]},
        'vitl': {'encoder': 'vitl', 'features': 256, 'out_channels': [256, 512, 1024, 1024]},
        'vitg': {'encoder': 'vitg', 'features': 384, 'out_channels': [1536, 1536, 1536, 1536]}
    }
    
    encoder = 'vits'
    dataset = 'hypersim'
    model = DepthAnythingV2(**model_configs[encoder])
    checkpoint_path = f'/path/to/checkpoints/depth_anything_v2_metric_{dataset}_{encoder}.pth'
    model.load_state_dict(torch.load(checkpoint_path, map_location='cpu'))
    model = model.to(DEVICE).eval()
    
    image_path = "/home/wasiq/Pictures/2m_Depth_Distance.jpeg"
    raw_img = cv2.imread(image_path)
    depth = model.infer_image(raw_img)
    print(f"The minimum depth is {depth.min()}")
    print(f"The maximum depth is {depth.max()}")
    
    **Observed behavior**
    xFormers not available printed (twice in logs)
    The minimum depth is 0.0
    The maximum depth is 0.0
    
    **Expected behavior**
    Non-constant depth map with varied values (for my test image the ground-truth near a marked point is ~2.0 m)
    depth.min() < depth.max() and meaningful spatial variation
    
    **Diagnostics I already tried**
    Confirmed checkpoint loads without crash (no obvious exceptions) but I haven't validated state_dict key names yet.
    Tried different encoder names (vits, vitb) consistent with checkpoint filenames.
    Ensured XFORMERS_DISABLE_MEMORY_EFFICIENT_ATTENTION=1 or uninstalled xFormers.
  4. heyoeyo commented on Dec 6, 2025

    @heyoeyo

    It looks like the relative depth model is being imported (from depth_anything_v2.dpt) instead of the metric depth model (from metric_depth.depth_anything_v2.dpt). They're nearly identical models, except for the very last layer where the relative depth model ends with a relu, which might be what's clamping all the values to 0.

  5. gururajbg commented on Jan 5, 2026

    @gururajbg

    Adding one more data point here based on local testing:

    Apart from importing the metric vs relative depth model, I also found that
    CPU inference can fail due to a device mismatch inside infer_image().

    When forcing device="cpu", the model weights are on CPU but the input tensor
    can still be created on CUDA, leading to either a runtime error or silent
    zero-depth output.

    Moving the input tensor to the model’s device inside infer_image() fixes
    CPU-only inference locally and produces non-constant depth.

    So the all-zero depth symptom seems to have multiple independent causes:

    1. importing the relative depth model instead of the metric one
    2. device mismatch during CPU inference
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions