Live data from Hacker News

Show HN: Running LLMs in one line of Python without Docker

lepton.ai

1–10 of 32 posts

Show HN: Running LLMs in one line of Python without Docker

#1
Hello Hacker News! We're Yangqing, Xiang and JJ from lepton.ai. We are building a platform to run any AI models as easy as writing local code, and to get your favorite models in minutes. It's like container for AI, but without the hassle of actually building a docker image.

We built and contributed to some of the world's most popular AI software - PyTorch 1.0, ONNX, Caffe, etcd, Kubernetes, etc. We also managed hundreds of thousands of computers in our previous jobs. And we found that the AI software stack is usually unnecessarily complex - and we want to change that.

Imagine if you are a developer who sees a good model on github, or HuggingFace. To make it a production ready service, the current solution usually requires you to build a docker image. But think about it - I have a few python code and a few python dependencies. That sounds like a huge overhead, right?

lepton.ai is really a pythonic way to free you from such difficulties. You write a simple python scaffold around your PyTorch / TensorFlow code, and lepton launches it as a full-fledged service callable via python, javascript, or any language that understands OpenAPI. We use containers under the hood, but you don't need to worry about all the infrastructure nuts and bolts.

One of the biggest challenge in AI is that it's really "all-stack": in addition to a plethora of models, AI applications usually involves GPUs, cloud infra, web services, DevOps, and SysOps. But we want you to focus on your job - and we take care of the rest "boring but essential" work.

We're really excited we get to show this to you all! Please let us know your thoughts and questions in the comments.

Show HN: Running LLMs in one line of Python without Docker
lepton.ai

Re: Show HN: Running LLMs in one line of Python without Docker

#4
To show some actual coding examples, We have made the python library open-source at https://github.com/leptonai/leptonai/. With it, launching a common HuggingFace model is as simple as a one liner. For example, if you have a GPU, Stable Diffusion XL is as simple as:

pip install -U leptonai

lep photon run -n sdxl -m hf:stabilityai/stable-diffusion-xl-base-1.0 --local

And you have a local OpenAPI server that runs it! Go to http://0.0.0.0:8080/docs, or use your favorite OpenAPI client.

We've been building AI API services using such tools ourselves. The easiest way to try out Lepton is to head to https://lepton.ai/playground and use our API service for popular models: Stable Diffusion, LLaMA, WhisperX, and other interesting showcases

We are proud of our performance. For example, we have probably the fastest LLaMA 7B and 70B model APIs, and it costs $0.8 to run 1 million tokens inference - we believe it's the most affordable one in the market. In addition, during the open beta phase, calling these services is free when you sign up for the Lepton AI platform.

Under the hood, we wrote a platform to allow you to run things easily on the cloud with ease. For example, if you find Pygmalion to be a great conversation model but you don't have a GPU, use lepton's Remote() capability to launch a service:

from leptonai import Remote

pygmalion = Remote("hf:PygmalionAI/pygmalion-2-7b", resource_shape="gpu.a10")

Wait a few minutes for the model to be downloaded and run, and you can now use it as if it were a standard python function:

print(pygmalion.run(inputs="Once upon a time", max_new_tokens=128))

If you are interested in the operational details, you can find fine-grained controls at https://dashboard.lepton.ai/ as a fully managed platform - we also support BYOC (bring your own compute) if you are an enterprise needing more autonomy over infrastructure.

Re: Show HN: Running LLMs in one line of Python without Docker

#5
Is not having to build a dockers image worth $100 a month? I do find server setup to be a pain but I think if I will use a model for a year, I can take the time(3-4 hours) to set it up. Only with constant switching of models would I use a service like this.

I never set up bigger models like LLAMA on servers. Other hacker news people can chime in.

Re: Show HN: Running LLMs in one line of Python without Docker

#7
post #5

Is not having to build a dockers image worth $100 a month? I do find server setup to be a pain but I think if I will use a model for a year, I can take the time(3-4 hours) to set it up. Only with constant switching of models would I use a service like this. I never set up bigger models like LLAMA on servers. Other hacker news people can chime in.

It's not only about "building a docker" but also maintaining multiple models, multiple environments and a lot of users. Imagine there is a group of engineers each needing to deploy their own models: one needs tensorflow 1.x, one needs tensorflow 2.x, one needs pytorch and one needs a very strange combination of dependencies. Trust me, things get complex very easily:

https://github.com/leptonai/examples/blob/main/advanced/whis...

I definitely agree that for a fixed use case, building a docker once and for all is probably the simplest and best approach. However, it quickly gets more complex and out of hand.

Also the basic plan is free for independent developers. You don't need to pay more than as if you were using EC2 instances, but with the platform convenience - we definitely hope it's worth it!

Re: Show HN: Running LLMs in one line of Python without Docker

#8

Looks interesting! Your llama 2 demo is unfortunately down: https://imgur.com/a/MLw6dAk

Great catch! Our cloud machine encountered a cuda error (the GPU fell off PCIe) - had to restart it. It's back to normal now.

All the more reason to have a managed version of services :)

Post reply on HN