Skip to content

feat(executorch): share one TensorRT engine between loads of the same program - #4778

Open
shoumikhin wants to merge 2 commits into
pytorch:mainfrom
shoumikhin:executorch-share-engines
Open

shoumikhin wants to merge 2 commits into
pytorch:mainfrom
shoumikhin:executorch-share-engines

Conversation

@shoumikhin

@shoumikhin shoumikhin commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

What is wrong today

A robot with two arms often runs the same policy twice in one process, from one .pte file. Each arm loads it:

method = Runtime.get().load_program(path).load_method("forward")

Each load deserializes a new TensorRT engine, which holds the weights, so the second arm pays for a second copy: about 800 MiB for SmolVLA on an 8 GB Jetson Orin Nano, where the CPU and GPU share one memory.

What this change does

Loads of the same engine bytes, device and weight streaming budget now share one engine. Each load keeps its own execution context and buffers; the last one frees the engine.

  • TensorRT cannot change the budget once a context exists, so the first load sets it. Another budget gets its own engine.
  • Engines are matched by a hash of their bytes and size, since the bytes are freed after loading. This trusts the program, as the README asks.
  • If two racing loads both deserialize and the second copy does not fit, that load shares the first engine instead of failing.
  • Setting the backend option use_shared_engines (on by default) to false turns sharing off for later loads. Only C++ can set it: the ExecuTorch Python bindings do not expose set_option.

The first commit only moves the budget code into a function.

Cost

Every load now hashes the engine bytes: about 0.15 s per GB on a Jetson AGX Thor, 0.25 s on the Orin Nano. Inference speed does not change.

Results

On the Orin Nano, a second load of the same SmolVLA program costs about 120 MiB instead of 800 MiB. Outputs are unchanged.

What was tested

  • New test_shared_engines: sharing, each part of the key, freeing, budgets, a failed load, no second deserialize on a match, the map lock while TensorRT logs, concurrent loads, two streams, dynamic shapes, the default and the option parser.
  • Of six deliberately broken backends, four fail a case: a key ignoring the bytes, a skipped first lookup, a failed load keeping the engine, and the map lock held during loading. A key ignoring the device needs a two-GPU case. A racing load that does not look again shows up only in a separate race under a memory cap.
  • The existing delegate tests give the same results before and after. The new test passed 20 runs in a row on a Jetson AGX Thor.

@meta-cla meta-cla Bot added the cla signed label Oct 4, 2026
@github-actions github-actions Bot added component: tests Issues re: Tests component: api [C++] Issues re: C++ API labels Oct 4, 2026
@github-actions
github-actions Bot requested a review from lanluo-nvidia October 4, 2026 13:54
@shoumikhin
shoumikhin force-pushed the executorch-share-engines branch from ba0b827 to c3f8053 Compare October 4, 2026 15:00
…unction

init applied the budget inline, between deserializing the engine and
creating its execution context. Move that block, unchanged, into
apply_weight_streaming_budget, called at the same point, so the next
change can run it once per engine rather than once per handle.

No behavior change.
@shoumikhin
shoumikhin force-pushed the executorch-share-engines branch 2 times, most recently from 56b0354 to 81987f8 Compare October 4, 2026 21:06
@shoumikhin
shoumikhin marked this pull request as ready for review October 4, 2026 21:07
… bytes

Loading one program twice in a process, for example once per robot arm,
deserialized its TensorRT engine twice, so the weights sat in GPU memory
twice. On an 8 GB Jetson Orin Nano, a second load of one robot policy cost
about 800 MiB.

Now loads of the same engine bytes, for the same device and the same
weight streaming budget, share one engine. Each load still creates its own
execution context, buffers and lock, so a second load costs about 120 MiB.
The last handle to go frees the engine.

The weight streaming budget is set before the engine's first context,
because TensorRT refuses to change it once a context exists. A load that
asks for a different budget gets its own engine.

The lock on the shared map is never held while TensorRT runs: TensorRT
logs while it loads, and under the Python bindings a log line waits for
the interpreter lock, so holding the map's lock across it can deadlock.
So two loads that race on the same bytes may both load an engine. The
second frees its copy and shares the first. If the second fails instead,
for example because the device has no room for a second copy, it looks
the engine up again and shares it if the first has published it by then.

Engines are matched by a hash of their bytes and their size, because the
bytes are freed after loading. This relies on programs being trusted
input, which the README already asks for. The key lives in a small
internal header, so a test can check it without a GPU.

Hashing makes every load slower: about 0.15 s per GB of engine on a
Jetson AGX Thor and 0.25 s per GB on the Orin Nano. Inference speed does
not change.

Setting the new use_shared_engines backend option, on by default, to
false turns sharing off for later loads. Only C++ can set it, since the
ExecuTorch Python bindings do not expose set_option.

The new test_shared_engines covers sharing, each part of the key, freeing,
budgets, a failed load, a match that must not deserialize again, the map
lock staying free while TensorRT logs, concurrent loads, two streams,
dynamic shapes, the default and the option parser. A case for the same
plan on two devices needs two GPUs and skips on one.
@shoumikhin
shoumikhin force-pushed the executorch-share-engines branch from 81987f8 to ef32a21 Compare October 5, 2026 05:45

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant