volumential.opcounters#

Explicit operation counters for the measurement drivers kept with the manuscripts (experiment E3).

The categories mirror the operation-count cost model that replaces wall-clock seconds as the primary cost currency (Section 6 of the paper):

  • SINGULAR_NODES — singular-quadrature node evaluations (Duffy triangle/cone nodes at which a radial profile or kernel is evaluated);

  • SMOOTH_NODES — smooth-rule (tensor Gauss-Legendre) node evaluations of the windowed remainder;

  • PROFILE_NODES — nodes passed through a windowed channel profile psi_m (each such node costs one special-function-class evaluation);

  • KERNEL_EVALS — kernel evaluations, keyed by the kernel’s special function;

  • SPECIAL_EVALS — special-function evaluations keyed by function (expn, erfc, exp, k0, hankel1, …);

  • RECOMBINATION_FLOPS — recombination fused multiply-adds and coefficient-recurrence steps of the assembly;

  • TABLE_ENTRIES — table entries built (one singular quadrature per stored entry).

Counting is strictly explicit: code that owns a loop calls add() (or OpCounters.add) with the executed element count; nothing is monkeypatched or inferred from timers. The module-level add() fans out to every OpCounters activated through the counting() context manager and is a no-op (one truthiness check) when none is active, so instrumented library code costs nothing outside instrumented runs.

The driver-level helpers at the bottom compute analytic counts from the executed configuration — the surviving-node counts of the requested quadrature rules, the symmetry-reduced entry set, the Duffy degeneracy predicates, and the FMM traversal — so drivers can emit measured-versus-analytic comparisons alongside their timing columns.

The last section adds the per-phase far-field counts of experiment E6: fmm_stage_operation_counts() prices the seven FMM stages of one solve from the traversal’s interaction lists and the wrangler’s expansion sizes, so that a driver pairing it with volumential.phase_profile can report a solve’s per-phase shares in operations and in seconds side by side.

volumential.opcounters.SINGULAR_NODES = 'singular_quadrature_nodes'#

Counter category for singular-quadrature node evaluations (Duffy triangle/cone nodes at which a radial profile or kernel is evaluated).

volumential.opcounters.SMOOTH_NODES = 'smooth_rule_nodes'#

Counter category for smooth-rule (tensor Gauss-Legendre) node evaluations of the windowed remainder.

volumential.opcounters.PROFILE_NODES = 'channel_profile_nodes'#

Counter category for nodes passed through a windowed channel profile psi_m; each costs one special-function-class evaluation.

volumential.opcounters.KERNEL_EVALS = 'kernel_evals'#

Counter category for kernel evaluations, keyed by the kernel’s special function.

volumential.opcounters.SPECIAL_EVALS = 'special_function_evals'#

Counter category for special-function evaluations keyed by function (expn, erfc, exp, k0, hankel1, …).

volumential.opcounters.RECOMBINATION_FLOPS = 'recombination_flops'#

Counter category for recombination fused multiply-adds and coefficient-recurrence steps of the assembly.

volumential.opcounters.TABLE_ENTRIES = 'table_entries_built'#

Counter category for table entries built (one singular quadrature per stored entry).

class volumential.opcounters.OpCounters[source]#

Bases: object

A bag of named operation counters, category -> label -> count.

add(category: str, label: str, count) → None[source]#

Add count to the counter named label within category.

labels(category: str) → dict[str, int][source]#

Per-label counts of one category (a copy).

total(category: str) → int[source]#

Sum of every label’s count in category, zero if it has none.

as_dict() → dict[str, dict[str, int]][source]#

A plain nested-dict copy of every counter, safe to serialize.

by_function(category: str) → str[source]#

Compact label:count listing, ‘;’-joined, sorted by label.

volumential.opcounters.counting(counters: OpCounters)[source]#

Activate counters for every add() call in the block.

volumential.opcounters.add(category: str, label: str, count) → None[source]#

Increment every active counter; no-op when none is active.

volumential.opcounters.direct_build_routing(table) → str[source]#

The DuffyRadial builder that actually produced table’s data.

One of the recorded routings (batched, scalar, scalar-adaptive, scalar-fallback), or unknown for a table whose payload predates the recording. Read the recorded routing rather than re-evaluating the routing predicate: _supports_batched_duffy_builder reports what was attempted, and a batched build that failed and fell back to the scalar per-entry builder has a different cost class and a different converged accuracy at fixed orders. A scalar-fallback here invalidates the batched node counts of the cost model for that table.

volumential.opcounters.direct_build_fallback_reason(table) → str[source]#

"<ExceptionType>: <message>" of a recorded scalar fallback, else the empty string.

volumential.opcounters.surviving_radial_node_count(radial_order, *, radial_rule='tanh-sinh-fast', mp_dps=50) → int[source]#

Surviving nodes of the requested radial rule (float64-saturating tanh-sinh nodes are masked by the builder, so the count is measured by executing the node builder, not assumed from the requested order).

volumential.opcounters.regular_node_count(dim, regular_order) → int[source]#

Regular (non-radial) node count of one Duffy region: the angular Gauss order in 2D, the two-axis tensor count in 3D — measured from the same node builder the quadratures use.

volumential.opcounters.batched_duffy_nodes_per_entry(dim, regular_order, radial_order, *, radial_rule='tanh-sinh-fast') → int[source]#

Node-template size of the batched direct builder per reduced entry: 2^d d! * regular nodes * surviving radial nodes (degenerate sign-octants keep their nodes in the generated code).

volumential.opcounters.reduced_entry_count(table) → int[source]#

Number of symmetry-reduced (quadratured) entries of a table.

volumential.opcounters.duffy_block_geometry(table) → dict[source]#

Counts of the grouped Duffy entry enumeration for one table geometry.

Mirrors the region-degeneracy predicates of the grouped channel builder (volumential.rke_table_assembly._duffy_channel_entry_values), which agree with the scalar builder’s collinear-triangle skip: 2D triangles are dropped when |det| < 1e-14 extent^2, 3D sign-octants when any edge length vanishes (each surviving octant contributes the d! axis-permutation cones).

Returns:

dict with n_reduced_entries, n_blocks (distinct (case, target) blocks), n_active_regions (non-degenerate triangles/cones summed over blocks — the grouped builders evaluate each region once per block), and entry_weighted_active_regions (the same regions summed per member entry — the scalar per-entry builder pays each region once per entry).

volumential.opcounters.grouped_duffy_singular_nodes(table, regular_order, radial_order, *, geometry=None) → int[source]#

Singular-quadrature nodes of one grouped (per-block) build over the reduced entry set — what one windowed/complex channel build evaluates.

volumential.opcounters.scalar_duffy_singular_nodes(table, regular_order, radial_order, *, geometry=None) → int[source]#

Singular-quadrature nodes of one scalar (per-entry) build over the reduced entry set — the legacy interpreted path pays every region once per member entry rather than once per block.

volumential.opcounters.nearfield_point_pairs_from_counts(*, target_boxes, neighbor_source_boxes_starts, neighbor_source_boxes_lists, box_target_counts_nonchild, box_source_counts_nonchild=None) → int[source]#

Near-field (target point, source point) pairs of one List 1 pass.

For every target box, its targets times the source points of its neighbor source boxes. box_source_counts_nonchild defaults to the target counts, which is the box-FMM near field’s own convention (a box’s sources are its target quadrature points). Pass an explicit source-count array to price a pass that runs over a different source set — the online split remainder does exactly that when it evaluates on an interpolated smooth quadrature instead of the base nodes.

volumential.opcounters.nearfield_point_pairs(queue, traversal) → int[source]#

Near-field (target point, source quadrature point) pairs of one List 1 pass: for every target box, its targets times the quadrature points of its neighbor source boxes. In the box-FMM near field, sources of a box are its target quadrature points (box_target_counts_nonchild serves both sides), matching the wrangler’s near-field kernel launch.

volumential.opcounters.FMM_FAR_FIELD_STAGES = ('form_multipoles', 'coarsen_multipoles', 'multipole_to_local', 'eval_multipoles', 'form_locals', 'refine_locals', 'eval_locals')#

FMM far-field stage names, matching volumential.phase_profile.FAR_FIELD_PHASES without the far_ prefix and the order in which drive_volume_fmm runs them.

volumential.opcounters.fmm_stage_operation_counts(*, box_levels, box_source_counts_nonchild, box_target_counts_nonchild, box_child_ids, source_boxes, source_parent_boxes, target_boxes, target_or_target_parent_boxes, from_sep_siblings_starts, from_sep_bigger_starts, from_sep_bigger_lists, sep_smaller_by_level, multipole_coeff_counts, local_coeff_counts) → dict[str, int][source]#

Operation counts of the seven FMM far-field stages of one solve.

Every count is a number of coefficient touches: one multiply-add against one expansion coefficient (or, for P2M/L2P/P2L, one source/target against one coefficient). This is the dense-translation structural model, instantiated from the executed traversal’s own interaction lists and the executed wrangler’s own per-level expansion sizes – nothing here is a constant, and nothing is a flop count of the generated OpenCL: a translation that sumpy accelerates (FFT-based or translation-class M2L, rotation-based operators) executes fewer machine operations than the dense count reported here, while still touching the same coefficients. Consumers must read the counts as a structural decomposition of the stage graph, not as a hardware work estimate.

The stage-by-stage rules follow the loops of sumpy.fmm.SumpyExpansionWrangler at the executed revision, so where the textbook model and the implementation disagree the implementation wins. Two such disagreements are reproduced here:

  • coarsen_multipoles (M2M) runs only for source levels nlevels-1 ... 3 (sumpy.fmm.SumpyExpansionWrangler .coarsen_multipoles), because no level-1 box is well separated from another; translations into levels 0 and 1 are never performed and are therefore not counted.

  • refine_locals (L2L) is driven by the target box list, one parent-to-child translation per target-or-target-parent box at level >= 1, not by an enumeration of children.

Parameters:
  • box_levels – tree.box_levels, host array.

  • box_source_counts_nonchild – tree.box_source_counts_nonchild.

  • box_target_counts_nonchild – tree.box_target_counts_nonchild.

  • box_child_ids – tree.box_child_ids, shape (2**dim, nboxes); a zero entry means “no such child” (boxtree convention).

  • source_boxes – traversal.source_boxes.

  • source_parent_boxes – traversal.source_parent_boxes.

  • target_boxes – traversal.target_boxes.

  • target_or_target_parent_boxes – traversal.target_or_target_parent_boxes.

  • from_sep_siblings_starts – List 2 CSR starts, indexed like target_or_target_parent_boxes.

  • from_sep_bigger_starts – List 4 CSR starts, indexed like target_or_target_parent_boxes.

  • from_sep_bigger_lists – List 4 CSR source-box lists.

  • sep_smaller_by_level – one (target_boxes, starts) pair per source level, from traversal.target_boxes_sep_smaller_by_source_level and traversal.from_sep_smaller_by_level[i].starts (List 3 far).

  • multipole_coeff_counts – multipole coefficients per tree level.

  • local_coeff_counts – local coefficients per tree level.

Returns:

a dict with one key per entry of FMM_FAR_FIELD_STAGES plus far_total.

volumential.opcounters.expansion_coefficient_counts(wrangler) → tuple[list[int], list[int]][source]#

Per-level (multipole, local) coefficient counts of a wrangler.

Read off the executed expansion classes at the executed per-level orders, so a kernel-specific expansion (the 2D Helmholtz/Yukawa Bessel expansions, say, at 2 * order + 1 coefficients) is priced as the run actually priced it rather than by a Taylor-order formula.

volumential.opcounters.fmm_stage_operation_counts_from_traversal(queue, traversal, wrangler) → dict[str, int][source]#

fmm_stage_operation_counts() for an executed traversal/wrangler.