volumential.phase_profile#

Per-phase wall-clock profiling for the measurement drivers kept with the manuscripts (experiment E6).

The operation-count cost model prices provisioning strategies (see volumential.opcounters). Experiment E6 asks the complementary question about an end-to-end solve: which phase of one solve owns which share of the work, in operations and in seconds. This module supplies the seconds half; volumential.opcounters.fmm_stage_operation_counts() supplies the operations half, phase for phase.

The phases of volumential.volume_fmm.drive_volume_fmm() are the FMM stage graph plus the two near-field phases:

phase name

stage

far_form_multipoles

P2M

far_coarsen_multipoles

M2M (upward pass)

nearfield_table_apply

List 1 through the base near-field table

split_correction

online split correction: series remainder P2P, the retained-channel table applies, and the smooth-source rebuild they need

far_multipole_to_local

M2L (List 2)

far_eval_multipoles

M2P (List 3 far)

far_form_locals

P2L (List 4 far)

far_refine_locals

L2L (downward pass)

far_eval_locals

L2P

Design constraints, in order of importance:

  1. Zero cost when inactive. phase() is a plain function that short-circuits on one ContextVar read and returns a shared do-nothing context manager when no profile is active, so the instrumented driver keeps its uninstrumented timings. Deliberately not a @contextmanager generator: building and entering one costs about 1.4 us, which drive_volume_fmm would pay eight or nine times per solve – eight versus nine, because only the online-split path enters split_correction, so the “inert” instrumentation would put a strategy-dependent bias into exactly the unprofiled solve means the break-even curves are built from. Nothing is monkeypatched and no timer runs unless a caller opted in.

  2. Activation follows the caller, not the process. The active profiles live in a ContextVar, so two threads that each profile a solve record only their own phases and drain only their own OpenCL queue. A process-global activation would silently merge them.

  3. Device work is attributed to the phase that launched it. OpenCL command queues are asynchronous, so a host-side timer around a kernel launch measures the launch, not the kernel. A profile therefore carries a sync callable (in practice queue.finish) which phase() invokes on entry and again before stopping the clock. This serializes the queue at every phase boundary, so a profiled solve is not a faithful measurement of an unprofiled solve’s total wall time: drivers must run profiled solves separately from the solves whose totals they report, and must report the profiled total alongside the shares so the perturbation is visible.

  4. Phases are disjoint by construction, not by assumption. The blocks in drive_volume_fmm do not nest. If a caller nests them anyway, the inner name is recorded in PhaseProfile.nested_names and the elapsed time is counted under both names; a consumer that finds nested_names non-empty must not read the shares as a partition.

Whatever the profiled phases do not cover – reordering sources and potentials, finalization, host bookkeeping – is the caller’s business to report as a residual (profiled solve total minus the sum of the phases).

volumential.phase_profile.FAR_FIELD_PHASES = ('far_form_multipoles', 'far_coarsen_multipoles', 'far_multipole_to_local', 'far_eval_multipoles', 'far_form_locals', 'far_refine_locals', 'far_eval_locals')#

FMM far-field stage phases, in the order drive_volume_fmm runs them.

volumential.phase_profile.NEAR_FIELD_PHASES = ('nearfield_table_apply', 'split_correction')#

Near-field phases of one solve, in the order drive_volume_fmm runs them. split_correction is entered only by the online split evaluator; a direct-table solve records zero seconds for it.

volumential.phase_profile.SOLVE_PHASES = ('far_form_multipoles', 'far_coarsen_multipoles', 'nearfield_table_apply', 'split_correction', 'far_multipole_to_local', 'far_eval_multipoles', 'far_form_locals', 'far_refine_locals', 'far_eval_locals')#

Every phase of one solve, in execution order.

class volumential.phase_profile.PhaseProfile(*, sync=None)[source]#

Bases: object

Accumulated wall seconds and entry counts, keyed by phase name.

Parameters:

sync – a callable invoked at every phase boundary to drain any asynchronous device queue (typically queue.finish), or None for pure host-side code.

sync() → None[source]#

Drain the device queue, if one was registered.

record(name: str, seconds: float, *, calls: int = 1) → None[source]#

Add seconds (and calls entries) to phase name.

property nested_names: frozenset[str]#

Phase names that were entered inside another phase.

Non-empty means the recorded seconds double count and are not a partition of anything.

names() → list[str][source]#

Recorded phase names, in first-recorded order.

seconds(name: str) → float[source]#

Total seconds recorded under name (0.0 if never entered).

calls(name: str) → int[source]#

Times name was entered (0 if never).

mean_seconds(name: str) → float[source]#

Seconds per entry of name (0.0 if never entered).

total_seconds() → float[source]#

Sum over all recorded phases.

as_dict() → dict[str, float][source]#

Copy of the accumulated seconds, keyed by phase name.

shares(*, denominator: float | None = None) → dict[str, float][source]#

Per-phase share of denominator (default: the phase total).

Returns an empty mapping when the denominator is zero, rather than inventing a share.

volumential.phase_profile.profiling(profile: PhaseProfile)[source]#

Activate profile for every phase() block in the body.

volumential.phase_profile.active() → bool[source]#

Whether any profile is currently collecting in this context.

volumential.phase_profile.phase(name: str)[source]#

Time the body under phase name; a no-op when nothing is active.

Returns a context manager rather than being one, so the inactive path is a ContextVar read and a shared-singleton return (see the private _InactivePhase for why that matters).