# Error running offline\_inference.py

**URL:** <https://forum.modular.com/t/error-running-offline-inference-py/1536>\
**Category:** MAX\
**Created:** [May 29, 2025, 12:12am UTC](https://forum.modular.com/t/error-running-offline-inference-py/1536 "2025-05-29T00:12:15Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![zetwhite](https://avatars.discourse-cdn.com/v4/letter/z/a6a055/32.png) [@zetwhite](https://forum.modular.com/u/zetwhite)\
**Post date:** [May 29, 2025, 12:12am UTC](https://forum.modular.com/t/error-running-offline-inference-py/1536/1 "2025-05-29T00:12:15Z")

</div>

Hi, I’m new to modular and I’m trying to run offline inference code from [Quickstart | Modular](https://docs.modular.com/max/get-started).

```python
from max.entrypoints.llm import LLM
from max.pipelines import PipelineConfig

def main():
    model_path = "modularai/Llama-3.1-8B-Instruct-GGUF"
    pipeline_config = PipelineConfig(model_path=model_path)
    llm = LLM(pipeline_config)

    prompts = [
        "In the beginning, there was",
        "I believe the meaning of life is",
        "The fastest way to learn python is",
    ]

    print("Generating responses...")
    responses = llm.generate(prompts, max_new_tokens=50)
    for i, (prompt, response) in enumerate(zip(prompts, responses)):
        print(f"========== Response {i} ==========")
        print(prompt + response)
        print()

if __name__ == " __main__":
    main()

```

But i got error with this log :

```bash
[2025-05-29 09:06:53] WARNING memory_estimation.py:142: Truncated model's default max_length from 131072 to 94767 to fit in memory.
[2025-05-29 09:06:53] INFO memory_estimation.py:190: 

	Estimated memory consumption:
	    Weights: 4.58 GiB
	    KVCache allocation: 23.14 GiB
	    Total estimated: 27.72 GiB used / 30.80 GiB free
	Auto-inferred max sequence length: 94767
	Auto-inferred max batch size: 1

Exception ignored in: <function LLM. __del__ at 0x734964a084c0>
Traceback (most recent call last):
  File "/home/zetwhite/.local/lib/python3.10/site-packages/max/entrypoints/llm.py", line 72, in __del__
    self._pc.set_canceled()
AttributeError: 'LLM' object has no attribute '_pc'
Traceback (most recent call last):
  File "/home/zetwhite/quickstart/offline.py", line 25, in <module>
    main()
  File "/home/zetwhite/quickstart/offline.py", line 8, in main
    llm = LLM(pipeline_config)
TypeError: LLM. __init__ () missing 1 required positional argument: 'pipeline_config'

```

I’m using ubuntu 22.04 / python 3.10 / modular == 25.3.0.  
What causes this error and how can I fix it?

---

<div class="post-metadata">

**Author:** ![zetwhite](https://avatars.discourse-cdn.com/v4/letter/z/a6a055/32.png) [@zetwhite](https://forum.modular.com/u/zetwhite)\
**Post date:** [May 29, 2025, 12:17am UTC](https://forum.modular.com/t/error-running-offline-inference-py/1536/2 "2025-05-29T00:17:30Z")

</div>

Ah, I found this post [Quickstart Docs](https://forum.modular.com/t/quickstart-docs/1501) and giving `setting` to LLM fixed my problem!

---

<div class="post-metadata">

**Author:** ![valon](https://sea1.discourse-cdn.com/flex001/user_avatar/forum.modular.com/valon/32/517_2.png) [@valon](https://forum.modular.com/u/valon)\
**Post date:** [May 29, 2025, 9:25pm UTC](https://forum.modular.com/t/error-running-offline-inference-py/1536/3 "2025-05-29T21:25:56Z")

</div>

Glad it was helpful.

Considering we’re seeing multiple people report the same problem I’m working to get the docs and code example updated.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/flex001/uploads/modular/original/1X/2751e0fbdc595a99718b216730957e9db4448cfd.jpeg) [@system](https://forum.modular.com/u/system)\
**Post date:** [November 25, 2025, 9:26pm UTC](https://forum.modular.com/t/error-running-offline-inference-py/1536/4 "2025-11-25T21:26:05Z")

</div>

This topic was automatically closed 180 days after the last reply. New replies are no longer allowed.
