Model Weight Logical Reshard Planner#
mooncake-reshard plans an address-free conversion between complete model
weight placements. It turns a source placement or committed logical Store
snapshot and a target placement into compact N-D transfer regions. It does not
inspect framework runtime objects or assign physical GPU addresses.
Inputs and Output#
The source is either a complete WeightPlacementManifest or a committed
WeightManifest snapshot. The target is a complete
WeightPlacementManifest. Both sides must identify the same resource,
revision, and weight generation.
The public APIs are:
plan_placement_transfer(source_placement, target_placement);plan_placement_transfer_to_local_target(source_placement, target_placement, target_participant_id);plan_stored_transfer_to_target_placement(source_manifest, target_placement).
Each API returns a LogicalTransferPlan. It contains only canonical tensor
descriptors, selected placement participants, and logical regions. It contains
no GPU address, endpoint, allocation range, lease, or backend handle.
N-D Regions#
Each TransferRegion represents one source/target N-D box overlap. It records
the overlap offset and shape, source and target base byte offsets, contiguous
inner_bytes, outer loop counts, and source/target byte strides.
The planner preserves a compact strided representation. It does not expand a
cross-dimension overlap into one operation per row or element. PlanningLimits
bounds the total number of regions and any later segment expansion fails closed
when it exceeds the configured limit.
Parallel Semantics#
TP changes logical boxes. The same overlap algorithm handles split, merge, and source/target sharding on different dimensions.
PP is explicit framework-provided tensor or layer ownership. Regions are grouped by source and target PP owner and optional pipeline stage; the planner does not infer ownership from a tensor name or layer-count formula.
EP is represented by a logical expert coordinate. Independent expert allocations remain independent logical fragments and are never packed or all-gathered by the planner.
DP does not change tensor geometry. A
ReplicatedAxis(kind="dp")uses a complete source replica. AnOwnershipAxis(kind="dp")routes each tensor through its declared owner and does not require every tensor on every DP rank.
All four axes are resolved by one logical-box plan, rather than by model-wide per-axis conversion passes.
Validation#
Placement construction validates the complete participant set, tensor descriptors, topology, and logical coverage before planning. Planning then fails closed when source and target tensor identity, dtype, shape, layout fingerprint, ownership, or coverage differ.
A WeightManifest source is retained as an immutable logical snapshot. Its
canonical identity and selected stored fragments are revalidated whenever a
logical plan is constructed or reconstructed. This proves that the plan still
refers to the same Store snapshot; it does not make Store persistence or
runtime loading part of this layer.
Coverage validation uses an ordered interval scan for 1-D inputs and a
coordinate-compressed sweep for 2-D inputs, both with O(N log N) behavior.
For 3-D and higher logical boxes, exact intersection remains supported under an
explicit pairwise-comparison budget; inputs that exceed it fail closed rather
than making validation work unbounded.
Boundary to the Next Phase#
This PR is intentionally limited to logical planning. A later runtime-binding
phase receives a LogicalTransferPlan and concrete
WeightRuntimeBindingManifest values, revalidates physical address bounds,
leases, generations, and alias scope, then produces an executor-facing bound
plan.
Transfer Engine lowering, DMA submission, Store persistence/lifecycle, and framework activation remain outside this logical planner. Framework adapters own model semantics and conversion into canonical manifests; Mooncake core does not infer those semantics from framework objects or parameter names.