feat(executorch): share one TensorRT engine between loads of the same program - #4778
Open
shoumikhin wants to merge 2 commits into
Open
shoumikhin wants to merge 2 commits into
shoumikhin wants to merge 2 commits into
Conversation
shoumikhin
force-pushed
the
executorch-share-engines
branch
from
October 4, 2026 15:00
ba0b827 to
c3f8053
Compare
…unction init applied the budget inline, between deserializing the engine and creating its execution context. Move that block, unchanged, into apply_weight_streaming_budget, called at the same point, so the next change can run it once per engine rather than once per handle. No behavior change.
shoumikhin
force-pushed
the
executorch-share-engines
branch
2 times, most recently
from
October 4, 2026 21:06
56b0354 to
81987f8
Compare
shoumikhin
marked this pull request as ready for review
October 4, 2026 21:07
… bytes Loading one program twice in a process, for example once per robot arm, deserialized its TensorRT engine twice, so the weights sat in GPU memory twice. On an 8 GB Jetson Orin Nano, a second load of one robot policy cost about 800 MiB. Now loads of the same engine bytes, for the same device and the same weight streaming budget, share one engine. Each load still creates its own execution context, buffers and lock, so a second load costs about 120 MiB. The last handle to go frees the engine. The weight streaming budget is set before the engine's first context, because TensorRT refuses to change it once a context exists. A load that asks for a different budget gets its own engine. The lock on the shared map is never held while TensorRT runs: TensorRT logs while it loads, and under the Python bindings a log line waits for the interpreter lock, so holding the map's lock across it can deadlock. So two loads that race on the same bytes may both load an engine. The second frees its copy and shares the first. If the second fails instead, for example because the device has no room for a second copy, it looks the engine up again and shares it if the first has published it by then. Engines are matched by a hash of their bytes and their size, because the bytes are freed after loading. This relies on programs being trusted input, which the README already asks for. The key lives in a small internal header, so a test can check it without a GPU. Hashing makes every load slower: about 0.15 s per GB of engine on a Jetson AGX Thor and 0.25 s per GB on the Orin Nano. Inference speed does not change. Setting the new use_shared_engines backend option, on by default, to false turns sharing off for later loads. Only C++ can set it, since the ExecuTorch Python bindings do not expose set_option. The new test_shared_engines covers sharing, each part of the key, freeing, budgets, a failed load, a match that must not deserialize again, the map lock staying free while TensorRT logs, concurrent loads, two streams, dynamic shapes, the default and the option parser. A case for the same plan on two devices needs two GPUs and skips on one.
shoumikhin
force-pushed
the
executorch-share-engines
branch
from
October 5, 2026 05:45
81987f8 to
ef32a21
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What is wrong today
A robot with two arms often runs the same policy twice in one process, from one
.ptefile. Each arm loads it:Each load deserializes a new TensorRT engine, which holds the weights, so the second arm pays for a second copy: about 800 MiB for SmolVLA on an 8 GB Jetson Orin Nano, where the CPU and GPU share one memory.
What this change does
Loads of the same engine bytes, device and weight streaming budget now share one engine. Each load keeps its own execution context and buffers; the last one frees the engine.
use_shared_engines(on by default) to false turns sharing off for later loads. Only C++ can set it: the ExecuTorch Python bindings do not exposeset_option.The first commit only moves the budget code into a function.
Cost
Every load now hashes the engine bytes: about 0.15 s per GB on a Jetson AGX Thor, 0.25 s on the Orin Nano. Inference speed does not change.
Results
On the Orin Nano, a second load of the same SmolVLA program costs about 120 MiB instead of 800 MiB. Outputs are unchanged.
What was tested
test_shared_engines: sharing, each part of the key, freeing, budgets, a failed load, no second deserialize on a match, the map lock while TensorRT logs, concurrent loads, two streams, dynamic shapes, the default and the option parser.