volumential.phase_profile#
Per-phase wall-clock profiling for the measurement drivers kept with the manuscripts (experiment E6).
The operation-count cost model prices provisioning strategies (see
volumential.opcounters). Experiment E6 asks the complementary
question about an end-to-end solve: which phase of one solve owns which
share of the work, in operations and in seconds. This module supplies the
seconds half; volumential.opcounters.fmm_stage_operation_counts()
supplies the operations half, phase for phase.
The phases of volumential.volume_fmm.drive_volume_fmm() are the FMM
stage graph plus the two near-field phases:
phase name |
stage |
|---|---|
|
P2M |
|
M2M (upward pass) |
|
List 1 through the base near-field table |
|
online split correction: series remainder P2P, the retained-channel table applies, and the smooth-source rebuild they need |
|
M2L (List 2) |
|
M2P (List 3 far) |
|
P2L (List 4 far) |
|
L2L (downward pass) |
|
L2P |
Design constraints, in order of importance:
Zero cost when inactive.
phase()is a plain function that short-circuits on oneContextVarread and returns a shared do-nothing context manager when no profile is active, so the instrumented driver keeps its uninstrumented timings. Deliberately not a@contextmanagergenerator: building and entering one costs about 1.4 us, whichdrive_volume_fmmwould pay eight or nine times per solve – eight versus nine, because only the online-split path enterssplit_correction, so the “inert” instrumentation would put a strategy-dependent bias into exactly the unprofiled solve means the break-even curves are built from. Nothing is monkeypatched and no timer runs unless a caller opted in.Activation follows the caller, not the process. The active profiles live in a
ContextVar, so two threads that each profile a solve record only their own phases and drain only their own OpenCL queue. A process-global activation would silently merge them.Device work is attributed to the phase that launched it. OpenCL command queues are asynchronous, so a host-side timer around a kernel launch measures the launch, not the kernel. A profile therefore carries a sync callable (in practice
queue.finish) whichphase()invokes on entry and again before stopping the clock. This serializes the queue at every phase boundary, so a profiled solve is not a faithful measurement of an unprofiled solve’s total wall time: drivers must run profiled solves separately from the solves whose totals they report, and must report the profiled total alongside the shares so the perturbation is visible.Phases are disjoint by construction, not by assumption. The blocks in
drive_volume_fmmdo not nest. If a caller nests them anyway, the inner name is recorded inPhaseProfile.nested_namesand the elapsed time is counted under both names; a consumer that findsnested_namesnon-empty must not read the shares as a partition.
Whatever the profiled phases do not cover – reordering sources and potentials, finalization, host bookkeeping – is the caller’s business to report as a residual (profiled solve total minus the sum of the phases).
- volumential.phase_profile.FAR_FIELD_PHASES = ('far_form_multipoles', 'far_coarsen_multipoles', 'far_multipole_to_local', 'far_eval_multipoles', 'far_form_locals', 'far_refine_locals', 'far_eval_locals')#
FMM far-field stage phases, in the order
drive_volume_fmmruns them.
- volumential.phase_profile.NEAR_FIELD_PHASES = ('nearfield_table_apply', 'split_correction')#
Near-field phases of one solve, in the order
drive_volume_fmmruns them.split_correctionis entered only by the online split evaluator; a direct-table solve records zero seconds for it.
- volumential.phase_profile.SOLVE_PHASES = ('far_form_multipoles', 'far_coarsen_multipoles', 'nearfield_table_apply', 'split_correction', 'far_multipole_to_local', 'far_eval_multipoles', 'far_form_locals', 'far_refine_locals', 'far_eval_locals')#
Every phase of one solve, in execution order.
- class volumential.phase_profile.PhaseProfile(*, sync=None)[source]#
Bases:
objectAccumulated wall seconds and entry counts, keyed by phase name.
- Parameters:
sync – a callable invoked at every phase boundary to drain any asynchronous device queue (typically
queue.finish), or None for pure host-side code.
- record(name: str, seconds: float, *, calls: int = 1) None[source]#
Add
seconds(andcallsentries) to phasename.
- property nested_names: frozenset[str]#
Phase names that were entered inside another phase.
Non-empty means the recorded seconds double count and are not a partition of anything.
Per-phase share of
denominator(default: the phase total).Returns an empty mapping when the denominator is zero, rather than inventing a share.
- volumential.phase_profile.profiling(profile: PhaseProfile)[source]#
Activate
profilefor everyphase()block in the body.
- volumential.phase_profile.active() bool[source]#
Whether any profile is currently collecting in this context.
- volumential.phase_profile.phase(name: str)[source]#
Time the body under phase
name; a no-op when nothing is active.Returns a context manager rather than being one, so the inactive path is a
ContextVarread and a shared-singleton return (see the private_InactivePhasefor why that matters).