Skip to content

Is possible an inference with just a RTX 3080 of 10GB? #25

Description

@davidmartinrius

Hello,

I know it is very little memory, but it is what I have by now.

By default, the demo code won't inference because of cuda out of memory. I tried to reduce the batch size of the inference to just 1, but is not enough.

Do you know a way to reduce the memory consumption running the inference?

I know that the best solution is to upgrade the GPU to a RTX 3090/4090/A6000, but before that I would like to try another way if possible.

Thank you!

David Martin Rius

Activity

  1. deepanwayx commented on Jul 4, 2023

    @deepanwayx
    Collaborator

    The required VRAM is around 13GB for full precision inference with a batch size of 1

    You can also try Colaboratory for inference: #10

  2. illtellyoulater commented on Jul 7, 2023

    @illtellyoulater

    @deepanwayx I suppose full inference precision is 32 bit, correct? If so, did you guys made any test to check whether 16 bit would still deliver good acceptable results?

  3. deepanwayx commented on Jul 10, 2023

    @deepanwayx
    Collaborator

    Yes, the full inference precision is 32-bit. We did not test with 16-bit inference.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions