Performance Models

Pruna Brings Ultra-Fast, Cost-Efficient Image and Video Generation to Eden AI

Victor Smith

DevRel @ Eden AI

Sara Han Díaz

DevRel Engineer

AI image and video stopped being a capability problem a while ago. You can generate a clip. The question is whether you can generate ten thousand of them this month without the bill or the latency becoming the reason the feature never ships.

That is the problem we work on. Pruna builds specialized performance models for image and video: near state-of-the-art quality at a fraction of the latency and cost. Speed and cost are not a side effect of our models, and neither is quality. That combination is the product.

Our models are now available through Eden AI, the European AI gateway, so teams can put them into production behind one key and one bill.

Speed comes from the optimization work

Our models grew out of Pruna's open-source optimization framework. Everything we learned there goes into our performance models, which is why they're faster, cheaper, smaller, and greener than other models.

That is why our image models are fast enough to sit inside an interactive product, and why our video models can afford a draft mode at all.

Where the money actually goes in video

AI video is billed per second of output, which means the expensive part is not the video you ship. It is the eleven versions you generated first and threw away.

A finished thirty-second clip is rarely thirty seconds of generation. It is a prompt you rewrote six times, a reference image you swapped twice, and a motion setting you only understood on the fourth try. Every one of those attempts was billed at full quality, for output nobody will ever watch.

Iteration does not need production quality only. You are checking composition, motion, and timing, and you can see all three in a lower-fidelity render. Our video models expose that as a draft flag, at roughly half the price of a full-quality render. Iterate in draft, switch it off for the take you keep.

In Eden AI, P-Video runs at $0.02 per second of output at 720p and $0.04 at 1080p. P-Video-2 is the higher-quality step up at $0.025 and $0.05, with draft mode at $0.015 and $0.03. P-Video-2-Pro is the top of the range, billed per second of returned video at $0.02 at 480p and $0.035 at 768p in speed mode, and $0.04 and $0.075 in quality mode. The useful pattern is draft on the cheaper model mode while you explore, full quality on the better model mode for the final render. Most teams do the opposite by default and pay for it.

The same lever on images

Image generation looks cheap per call, which is exactly why nobody optimizes it, and why it becomes one of the largest lines on the bill the moment a feature reaches real users.

At 10,000 requests a day, our image editing model comes out at roughly a third of the cost of a comparable model such as Flux Kontext Dev. Nothing about the output changes. The savings come from the optimization work underneath.

In production, through one API

Our models sit on Eden AI alongside the rest of its catalog, with one key, one bill, and the same asynchronous pattern for every video model. Calling P-Video-2 is one request:

curl -X POST <https://api.edenai.run/v3/universal-ai/async> \
  -H "Authorization: Bearer $EDENAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "video/generation_async/pruna/p-video-2",
    "input": {
      "text": "a slow aerial shot over a coastal city at sunrise",
      "duration": 5,
      "dimension": "1280x720"
    }
  }'
curl -X POST <https://api.edenai.run/v3/universal-ai/async> \
  -H "Authorization: Bearer $EDENAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "video/generation_async/pruna/p-video-2",
    "input": {
      "text": "a slow aerial shot over a coastal city at sunrise",
      "duration": 5,
      "dimension": "1280x720"
    }
  }'
curl -X POST <https://api.edenai.run/v3/universal-ai/async> \
  -H "Authorization: Bearer $EDENAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "video/generation_async/pruna/p-video-2",
    "input": {
      "text": "a slow aerial shot over a coastal city at sunrise",
      "duration": 5,
      "dimension": "1280x720"
    }
  }'

Changing which model you are testing is a string in the request, not an integration. That turns a real cost comparison into an afternoon's work instead of a procurement cycle per provider.

Give a text model the ability to generate video

Most language models cannot generate video. They can write you a prompt for one, and that is usually where it stops.

Eden AI publishes its expert models as tools on a hosted MCP server, so any function-calling model can call video generation itself instead of handing the job back to you. Point a client at https://mcp.edenai.run/mcp with your Eden AI key, and video_generation_async appears in the model's tool list with video/generation_async/pruna/p-video-2 among its options. Nothing to deploy, and the call is billed to the same key as the rest of your usage.

The loop is worth picturing. An agent writes a prompt, fires the generation, gets a job back, polls check_job, and returns the finished clip. Because P-Video-2 exposes draft mode, the agent can run its own iteration pass at half price and only commit to a full-quality render once the composition holds. That is the discard problem above, handled by the agent rather than by a person.

Built in Europe, available everywhere

Pruna is based in Munich and Paris, and Eden AI is a French company. For a European team, that means the two companies you are contracting with, and the people you can get on a call, are on your side of the Atlantic. That is a shorter conversation than it sounds, and it usually happens long before anyone runs a benchmark.

The models themselves are available worldwide through Eden AI's API, with no separate account to open wherever your team sits.

Where to start

Measure your discard rate first. If you keep one generation in five, draft mode is the single biggest saving available to you, and it costs nothing to adopt. If you keep four in five, your problem is somewhere else, and cheaper renders will not fix it.

Curious what Pruna can do for your models?

Whether you're running GenAI in production or exploring what's possible, Pruna makes it easier to move fast and stay efficient.

Curious what Pruna can do for your models?

Whether you're running GenAI in production or exploring what's possible, Pruna makes it easier to move fast and stay efficient.

Curious what Pruna can do for your models?

Whether you're running GenAI in production or exploring what's possible, Pruna makes it easier to move fast and stay efficient.