Skip to content

guidellm.benchmark.outputs

Output formatters for benchmark results.

Provides output formatter implementations that transform benchmark reports into various file formats including JSON, CSV, HTML, and console display. All formatters extend the base GenerativeBenchmarkerOutput interface, enabling dynamic resolution and flexible output configuration for benchmark result persistence and analysis.

GenerativeBenchmarkerCSV

Bases: GenerativeBenchmarkerOutput

CSV output formatter for benchmark results.

Exports comprehensive benchmark data to CSV format with multi-row headers organizing metrics into categories including run information, timing, request counts, latency, throughput, modality-specific data, and scheduler state. Each benchmark run becomes a row with statistical distributions represented as mean, median, standard deviation, and percentiles.

Attributes:

Name Type Description
DEFAULT_FILE str

Default filename for CSV output

Source code in src/guidellm/benchmark/outputs/csv.py
 73
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
631
632
633
634
635
636
637
638
639
640
641
642
643
644
645
646
647
648
649
650
651
652
653
654
655
656
657
658
659
660
661
662
663
664
665
666
667
668
669
670
671
672
673
674
675
676
677
678
679
680
681
682
683
684
685
686
687
688
689
690
691
692
693
694
695
696
697
698
699
700
701
702
703
704
705
706
707
708
709
710
711
712
713
714
715
716
717
718
719
720
721
722
723
724
725
726
727
728
729
730
731
732
733
734
735
736
737
738
739
740
741
742
743
744
745
746
747
748
749
750
751
752
753
754
755
756
757
758
759
760
761
762
763
764
765
766
767
768
769
770
771
772
773
774
775
776
777
778
779
780
781
782
783
784
785
786
787
788
789
790
791
792
793
794
795
796
797
798
799
800
801
802
803
804
805
806
807
808
809
810
811
812
813
814
815
816
817
818
819
820
821
822
823
824
825
826
827
828
829
830
831
832
833
834
835
836
837
838
839
840
841
842
843
844
845
846
847
848
849
850
851
852
853
854
855
856
857
858
859
860
861
862
863
864
865
866
867
868
869
870
871
872
873
874
875
876
877
878
879
880
881
882
883
884
885
886
887
888
889
890
891
892
893
894
895
896
897
898
899
900
901
902
903
904
905
906
907
908
909
910
911
912
913
914
915
916
917
918
919
920
921
922
923
924
925
926
927
928
929
930
931
932
933
934
935
936
937
938
939
940
941
942
943
944
945
946
947
@GenerativeBenchmarkerOutput.register("csv")
class GenerativeBenchmarkerCSV(GenerativeBenchmarkerOutput):
    """
    CSV output formatter for benchmark results.

    Exports comprehensive benchmark data to CSV format with multi-row headers
    organizing metrics into categories including run information, timing, request
    counts, latency, throughput, modality-specific data, and scheduler state. Each
    benchmark run becomes a row with statistical distributions represented as
    mean, median, standard deviation, and percentiles.

    :cvar DEFAULT_FILE: Default filename for CSV output
    """

    DEFAULT_FILE: ClassVar[str] = "benchmarks.csv"

    @classmethod
    def from_args(cls, args: BenchmarkOutputArgs) -> GenerativeBenchmarkerCSV:
        """
        Create a CSV output formatter from output arguments.

        :param args: Output configuration with path
        :return: Configured CSV output formatter
        """
        if not isinstance(args, CSVBenchmarkOutputArgs):
            raise ValueError(f"Expected CSVBenchmarkOutputArgs, got {type(args)}")

        return cls(output_path=args.path)

    output_path: Path = Field(
        default_factory=Path.cwd,
        description=(
            "Path where the CSV file will be saved, defaults to current directory"
        ),
    )

    async def finalize(self, report: GenerativeBenchmarksReport) -> Path:
        """
        Save the benchmark report as a CSV file.

        :param report: The completed benchmark report
        :return: Path to the saved CSV file
        """
        output_path = self.output_path
        if output_path.is_dir():
            output_path = output_path / GenerativeBenchmarkerCSV.DEFAULT_FILE
        output_path.parent.mkdir(parents=True, exist_ok=True)

        with output_path.open("w", newline="") as file:
            writer = csv.writer(file)

            all_headers: list[list[list[str]]] = []
            all_values: list[list[str | int | float]] = []

            for benchmark in report.benchmarks:
                benchmark_headers: list[list[str]] = []
                benchmark_values: list[str | int | float] = []

                self._add_run_info(benchmark, benchmark_headers, benchmark_values)
                self._add_benchmark_info(benchmark, benchmark_headers, benchmark_values)
                self._add_timing_info(benchmark, benchmark_headers, benchmark_values)
                self._add_request_counts(benchmark, benchmark_headers, benchmark_values)
                self._add_request_latency_metrics(
                    benchmark, benchmark_headers, benchmark_values
                )
                self._add_server_throughput_metrics(
                    benchmark, benchmark_headers, benchmark_values
                )
                for modality_name in MODALITY_METRICS:
                    self._add_modality_metrics(
                        benchmark,
                        modality_name,
                        benchmark_headers,
                        benchmark_values,
                    )
                self._add_scheduler_info(benchmark, benchmark_headers, benchmark_values)
                self._add_runtime_info(report, benchmark_headers, benchmark_values)
                self._add_interval_columns(
                    benchmark, benchmark_headers, benchmark_values
                )

                all_headers.append(benchmark_headers)
                all_values.append(benchmark_values)

            headers, data_rows = self._align_columns(all_headers, all_values)

            self._write_multirow_header(writer, headers)
            for row in data_rows:
                writer.writerow(row)

        return output_path

    def _add_interval_columns(
        self,
        benchmark: GenerativeBenchmark,
        headers: list[list[str]],
        values: list[str | int | float],
    ) -> None:
        """
        Add confidence interval columns for the metrics that carry them.

        Written after every other column so that existing column positions are
        unchanged. Only metrics recorded once per request carry intervals, so
        only those are listed, under the same labels the row already uses for
        them.

        :param benchmark: Benchmark data to extract intervals from
        :param headers: List of header hierarchies to append to
        :param values: List of values to append to
        """
        self._add_field(
            headers,
            values,
            "Measurement Uncertainty",
            "Confidence Level",
            "" if benchmark.config.confidence is None else benchmark.config.confidence,
        )

        metrics = benchmark.metrics
        interval_metrics: list[tuple[StatusDistributionSummary | None, str, str]] = [
            (metrics.request_latency, "Request Latency", "Sec"),
            (metrics.request_dispatch_delay, "Dispatch Delay", "Sec"),
            (metrics.request_scheduled_latency, "Scheduled Latency", "Sec"),
            (
                metrics.request_streaming_iterations_count,
                "Streaming Iterations",
                "Count",
            ),
            (metrics.time_to_first_token_ms, "Time to First Token", "ms"),
            (
                metrics.time_to_first_output_token_ms,
                "Time to First Output Token",
                "ms",
            ),
            (metrics.time_to_last_round_trip_ms, "Time To Last Round Trip", "ms"),
            (metrics.avg_round_trip_time_ms, "Avg Round Trip Time", "ms"),
            (metrics.prompt_token_count, "Token Metrics", "Input Tokens"),
            (metrics.output_token_count, "Token Metrics", "Output Tokens"),
            (metrics.total_token_count, "Token Metrics", "Total Tokens"),
        ]

        for metric, group, units in interval_metrics:
            if metric is None:
                continue
            for status, dist in (
                ("Successful", metric.successful),
                ("Incomplete", metric.incomplete),
                ("Errored", metric.errored),
            ):
                # Skip the statuses the metric columns skip, so each interval
                # column has a matching set of statistics earlier in the row.
                if dist.total_sum == 0.0:
                    continue
                headers.append([group, f"{status} {units}", "Mean CI"])
                values.append(
                    ""
                    if dist.mean_ci is None
                    else f"[{dist.mean_ci.lower}, {dist.mean_ci.upper}]"
                )
                headers.append([group, f"{status} {units}", "Percentile CIs"])
                values.append(self._format_percentile_intervals(dist))

    @staticmethod
    def _format_percentile_intervals(dist: DistributionSummary) -> str:
        """
        Render the percentile intervals as one JSON object keyed by percentile.

        Every percentile is present, so one without an interval reads as null
        rather than as missing.

        :param dist: Distribution summary to read the intervals from
        :return: JSON object of percentile to [lower, upper] or null, or an
            empty string when no intervals were estimated
        """
        if dist.percentile_cis is None:
            return ""

        return json.dumps(
            {
                name: None if bounds is None else [bounds["lower"], bounds["upper"]]
                for name, bounds in dist.percentile_cis.model_dump().items()
            },
            separators=(",", ":"),
        )

    @staticmethod
    def _align_columns(
        all_headers: list[list[list[str]]],
        all_values: list[list[str | int | float]],
    ) -> tuple[list[list[str]], list[list[str | int | float]]]:
        """
        Align columns across multiple benchmarks that may have different column sets.

        Builds a unified header list from all benchmarks (preserving first-seen order)
        and pads each row with empty strings for columns it doesn't have.

        :param all_headers: Per-benchmark list of column header hierarchies
        :param all_values: Per-benchmark list of column values
        :return: Tuple of (unified headers, aligned data rows)
        """
        ordered_headers: dict[tuple[str, ...], None] = {}
        row_maps: list[dict[tuple[str, ...], str | int | float]] = []

        for benchmark_headers, benchmark_values in zip(
            all_headers, all_values, strict=True
        ):
            row_map: dict[tuple[str, ...], str | int | float] = {}
            for header_parts, value in zip(
                benchmark_headers, benchmark_values, strict=False
            ):
                header_key = tuple(header_parts)
                row_map[header_key] = value
                if header_key not in ordered_headers:
                    ordered_headers[header_key] = None
            row_maps.append(row_map)

        header_keys = list(ordered_headers.keys())
        headers = [list(k) for k in header_keys]
        data_rows: list[list[str | int | float]] = [
            [row_map.get(k, "") for k in header_keys] for row_map in row_maps
        ]
        return headers, data_rows

    def _write_multirow_header(self, writer: Any, headers: list[list[str]]) -> None:
        """
        Write multi-row header to CSV for hierarchical metric organization.

        :param writer: CSV writer instance
        :param headers: List of column header hierarchies as string lists
        """
        max_rows = max((len(col) for col in headers), default=0)
        for row_idx in range(max_rows):
            row = [col[row_idx] if row_idx < len(col) else "" for col in headers]
            writer.writerow(row)

    def _add_field(
        self,
        headers: list[list[str]],
        values: list[str | int | float],
        group: str,
        field_name: str,
        value: Any,
        units: str = "",
    ) -> None:
        """
        Add a single field to headers and values lists.

        :param headers: List of header hierarchies to append to
        :param values: List of values to append to
        :param group: Top-level category for the field
        :param field_name: Name of the field
        :param value: Value for the field
        :param units: Optional units for the field
        """
        headers.append([group, field_name, units])
        values.append(value)

    def _add_runtime_info(
        self,
        report: GenerativeBenchmarksReport,
        headers: list[list[str]],
        values: list[str | int | float],
    ) -> None:
        """
        Add global metadata and environment information.

        :param report: Benchmark report to extract global info from
        :param headers: List of header hierarchies to append to
        :param values: List of values to append to
        """
        self._add_field(
            headers,
            values,
            "Runtime Info",
            "Metadata",
            report.metadata.model_dump_json(),
        )
        self._add_field(
            headers,
            values,
            "Runtime Info",
            "Arguments",
            report.config.model_dump_json(),
        )

    def _add_run_info(
        self,
        benchmark: GenerativeBenchmark,
        headers: list[list[str]],
        values: list[str | int | float],
    ) -> None:
        """
        Add overall run identification and configuration information.

        :param benchmark: Benchmark data to extract run info from
        :param headers: List of header hierarchies to append to
        :param values: List of values to append to
        """
        self._add_field(headers, values, "Run Info", "Run ID", benchmark.config.run_id)
        self._add_field(
            headers, values, "Run Info", "Run Index", benchmark.config.run_index
        )
        self._add_field(
            headers,
            values,
            "Run Info",
            "Profile",
            json.dumps(benchmark.config.profile),
        )
        self._add_field(
            headers,
            values,
            "Run Info",
            "Requests",
            json.dumps(benchmark.config.requests),
        )
        self._add_field(
            headers, values, "Run Info", "Backend", json.dumps(benchmark.config.backend)
        )
        self._add_field(
            headers,
            values,
            "Run Info",
            "Environment",
            json.dumps(benchmark.config.environment),
        )

    def _add_benchmark_info(
        self,
        benchmark: GenerativeBenchmark,
        headers: list[list[str]],
        values: list[str | int | float],
    ) -> None:
        """
        Add individual benchmark configuration details.

        :param benchmark: Benchmark data to extract configuration from
        :param headers: List of header hierarchies to append to
        :param values: List of values to append to
        """
        self._add_field(headers, values, "Benchmark", "Type", benchmark.type_)
        self._add_field(headers, values, "Benchmark", "ID", benchmark.config.id_)
        self._add_field(
            headers, values, "Benchmark", "Strategy", benchmark.config.strategy.type_
        )
        self._add_field(
            headers,
            values,
            "Benchmark",
            "Constraints",
            json.dumps(benchmark.config.constraints),
        )

    def _add_timing_info(
        self,
        benchmark: GenerativeBenchmark,
        headers: list[list[str]],
        values: list[str | int | float],
    ) -> None:
        """
        Add timing information including start, end, duration, warmup, and cooldown.

        :param benchmark: Benchmark data to extract timing from
        :param headers: List of header hierarchies to append to
        :param values: List of values to append to
        """
        timing_fields: list[tuple[str, Any]] = [
            ("Start Time", benchmark.scheduler_metrics.start_time),
            ("Request Start Time", benchmark.scheduler_metrics.request_start_time),
            ("Measure Start Time", benchmark.scheduler_metrics.measure_start_time),
            ("Measure End Time", benchmark.scheduler_metrics.measure_end_time),
            ("Request End Time", benchmark.scheduler_metrics.request_end_time),
            ("End Time", benchmark.scheduler_metrics.end_time),
        ]
        for field_name, timestamp in timing_fields:
            self._add_field(
                headers,
                values,
                "Timings",
                field_name,
                safe_format_timestamp(timestamp, TIMESTAMP_FORMAT),
            )

        duration_fields: list[tuple[str, float | str]] = [
            ("Duration", benchmark.duration),
            ("Warmup", benchmark.warmup_duration),
            ("Cooldown", benchmark.cooldown_duration),
        ]
        for field_name, duration_value in duration_fields:
            self._add_field(
                headers, values, "Timings", field_name, duration_value, "Sec"
            )

    def _add_request_counts(
        self,
        benchmark: GenerativeBenchmark,
        headers: list[list[str]],
        values: list[str | int | float],
    ) -> None:
        """
        Add request count totals by status.

        :param benchmark: Benchmark data to extract request counts from
        :param headers: List of header hierarchies to append to
        :param values: List of values to append to
        """
        for status in ["successful", "incomplete", "errored", "total"]:
            self._add_field(
                headers,
                values,
                "Request Counts",
                status.capitalize(),
                getattr(benchmark.metrics.request_totals, status),
            )

    def _add_request_latency_metrics(
        self,
        benchmark: GenerativeBenchmark,
        headers: list[list[str]],
        values: list[str | int | float],
    ) -> None:
        """
        Add request latency and streaming metrics.

        :param benchmark: Benchmark data to extract latency metrics from
        :param headers: List of header hierarchies to append to
        :param values: List of values to append to
        """
        self._add_stats_for_metric(
            headers, values, benchmark.metrics.request_latency, "Request Latency", "Sec"
        )
        # None for strategies without an arrival schedule; emit no columns.
        if benchmark.metrics.request_dispatch_delay is not None:
            self._add_stats_for_metric(
                headers,
                values,
                benchmark.metrics.request_dispatch_delay,
                "Dispatch Delay",
                "Sec",
            )
        if benchmark.metrics.turn_predecessor_delay is not None:
            self._add_stats_for_metric(
                headers,
                values,
                benchmark.metrics.turn_predecessor_delay,
                "Turn Predecessor Delay",
                "Sec",
            )
        if benchmark.metrics.turn_scheduling_delay is not None:
            self._add_stats_for_metric(
                headers,
                values,
                benchmark.metrics.turn_scheduling_delay,
                "Turn Scheduling Delay",
                "Sec",
            )
        if benchmark.metrics.request_scheduled_latency is not None:
            self._add_stats_for_metric(
                headers,
                values,
                benchmark.metrics.request_scheduled_latency,
                "Scheduled Latency",
                "Sec",
            )
        self._add_stats_for_metric(
            headers,
            values,
            benchmark.metrics.request_streaming_iterations_count,
            "Streaming Iterations",
            "Count",
        )
        self._add_stats_for_metric(
            headers,
            values,
            benchmark.metrics.time_to_first_token_ms,
            "Time to First Token",
            "ms",
        )
        self._add_stats_for_metric(
            headers,
            values,
            benchmark.metrics.time_to_first_output_token_ms,
            "Time to First Output Token",
            "ms",
        )
        self._add_stats_for_metric(
            headers,
            values,
            benchmark.metrics.time_per_output_token_ms,
            "Time per Output Token",
            "ms",
        )
        self._add_stats_for_metric(
            headers,
            values,
            benchmark.metrics.inter_token_latency_ms,
            "Inter Token Latency",
            "ms",
        )
        self._add_stats_for_metric(
            headers,
            values,
            benchmark.metrics.time_to_last_round_trip_ms,
            "Time To Last Round Trip",
            "ms",
        )
        self._add_stats_for_metric(
            headers,
            values,
            benchmark.metrics.avg_round_trip_time_ms,
            "Avg Round Trip Time",
            "ms",
        )

    def _add_server_throughput_metrics(
        self,
        benchmark: GenerativeBenchmark,
        headers: list[list[str]],
        values: list[str | int | float],
    ) -> None:
        """
        Add server throughput metrics including requests, tokens, and concurrency.

        :param benchmark: Benchmark data to extract throughput metrics from
        :param headers: List of header hierarchies to append to
        :param values: List of values to append to
        """
        self._add_stats_for_metric(
            headers,
            values,
            benchmark.metrics.requests_per_second,
            "Server Throughput",
            "Requests/Sec",
        )
        self._add_stats_for_metric(
            headers,
            values,
            benchmark.metrics.request_concurrency,
            "Server Throughput",
            "Concurrency",
        )
        # Emitted whenever objectives were configured, so a workload that
        # cannot evaluate them still reports attainment as empty rather than
        # dropping the columns and reading as though none were set. Every
        # benchmark in a run shares one objective set, so the columns stay
        # aligned across rows.
        if benchmark.config.slo is not None:
            # Written directly rather than through _add_stats_for_metric,
            # which drops any status whose total is 0.0. A run that conforms to
            # nothing would otherwise lose these columns from the CSV while the
            # console and JSON still report 0.0.
            goodput = benchmark.metrics.request_goodput
            headers.append(["Server Throughput", "Successful Goodput/Sec", "mean"])
            values.append("" if goodput is None else goodput.successful.mean)
            headers.append(["Server Throughput", "SLO Attainment", ""])
            attainment = benchmark.metrics.slo_attainment
            values.append("" if attainment is None else attainment)
        self._add_stats_for_metric(
            headers,
            values,
            benchmark.metrics.prompt_token_count,
            "Token Metrics",
            "Input Tokens",
        )
        self._add_stats_for_metric(
            headers,
            values,
            benchmark.metrics.output_token_count,
            "Token Metrics",
            "Output Tokens",
        )
        self._add_stats_for_metric(
            headers,
            values,
            benchmark.metrics.total_token_count,
            "Token Metrics",
            "Total Tokens",
        )
        self._add_stats_for_metric(
            headers,
            values,
            benchmark.metrics.prompt_tokens_per_second,
            "Token Throughput",
            "Input Tokens/Sec",
        )
        self._add_stats_for_metric(
            headers,
            values,
            benchmark.metrics.output_tokens_per_second,
            "Token Throughput",
            "Output Tokens/Sec",
        )
        self._add_stats_for_metric(
            headers,
            values,
            benchmark.metrics.tokens_per_second,
            "Token Throughput",
            "Total Tokens/Sec",
        )
        self._add_stats_for_metric(
            headers,
            values,
            benchmark.metrics.output_tokens_per_iteration,
            "Token Streaming",
            "Output Tokens/Iter",
        )
        self._add_stats_for_metric(
            headers,
            values,
            benchmark.metrics.iter_tokens_per_iteration,
            "Token Streaming",
            "Iter Tokens/Iter",
        )

    def _add_modality_metrics(
        self,
        benchmark: GenerativeBenchmark,
        modality: str,
        headers: list[list[str]],
        values: list[str | int | float],
    ) -> None:
        """
        Add modality-specific metrics for text, image, video, audio, or tool calls.

        :param benchmark: Benchmark data to extract modality metrics from
        :param modality: Type of modality to extract metrics for
        :param headers: List of header hierarchies to append to
        :param values: List of values to append to
        """
        modality_summary = getattr(benchmark.metrics, modality)
        metric_definitions = MODALITY_METRICS[modality]

        for metric_name, display_name in metric_definitions:
            metric_obj = getattr(modality_summary, metric_name, None)
            if metric_obj is None:
                continue

            for io_type in ["input", "output", "total"]:
                dist_summary = getattr(metric_obj, io_type, None)
                if dist_summary is None:
                    continue

                if not self._has_distribution_data(dist_summary):
                    continue

                self._add_stats_for_metric(
                    headers,
                    values,
                    dist_summary,
                    f"{modality.replace('_', ' ').title()} {display_name}",
                    io_type.capitalize(),
                )

    def _has_distribution_data(self, dist_summary: StatusDistributionSummary) -> bool:
        """
        Check if distribution summary contains any data.

        Uses ``count > 0`` rather than ``total_sum > 0`` so that
        all-zero distributions (e.g. errored tool-call requests) are
        still recognised as having data.

        :param dist_summary: Distribution summary to check
        :return: True if summary contains data, False otherwise
        """
        return any(
            getattr(dist_summary, status, None) is not None
            and getattr(dist_summary, status).count > 0
            for status in ["successful", "incomplete", "errored"]
        )

    def _add_scheduler_info(
        self,
        benchmark: GenerativeBenchmark,
        headers: list[list[str]],
        values: list[str | int | float],
    ) -> None:
        """
        Add scheduler state and performance information.

        :param benchmark: Benchmark data to extract scheduler info from
        :param headers: List of header hierarchies to append to
        :param values: List of values to append to
        """
        self._add_scheduler_state(benchmark, headers, values)
        self._add_scheduler_metrics(benchmark, headers, values)

    def _add_scheduler_state(
        self,
        benchmark: GenerativeBenchmark,
        headers: list[list[str]],
        values: list[str | int | float],
    ) -> None:
        """
        Add scheduler state information including request counts and timing.

        :param benchmark: Benchmark data to extract scheduler state from
        :param headers: List of header hierarchies to append to
        :param values: List of values to append to
        """
        state = benchmark.scheduler_state

        state_fields: list[tuple[str, Any]] = [
            ("Node ID", state.node_id),
            ("Num Processes", state.num_processes),
            ("Created Requests", state.created_requests),
            ("Processed Requests", state.processed_requests),
            ("Successful Requests", state.successful_requests),
            ("Errored Requests", state.errored_requests),
            ("Cancelled Requests", state.cancelled_requests),
        ]

        for field_name, value in state_fields:
            self._add_field(headers, values, "Scheduler State", field_name, value)

        if state.end_queuing_time:
            self._add_field(
                headers,
                values,
                "Scheduler State",
                "End Queuing Time",
                safe_format_timestamp(state.end_queuing_time, TIMESTAMP_FORMAT),
            )
            end_queuing_constraints_dict = {
                key: constraint.model_dump()
                for key, constraint in state.end_queuing_constraints.items()
            }
            self._add_field(
                headers,
                values,
                "Scheduler State",
                "End Queuing Constraints",
                json.dumps(end_queuing_constraints_dict),
            )

        if state.end_processing_time:
            self._add_field(
                headers,
                values,
                "Scheduler State",
                "End Processing Time",
                safe_format_timestamp(state.end_processing_time, TIMESTAMP_FORMAT),
            )
            end_processing_constraints_dict = {
                key: constraint.model_dump()
                for key, constraint in state.end_processing_constraints.items()
            }
            self._add_field(
                headers,
                values,
                "Scheduler State",
                "End Processing Constraints",
                json.dumps(end_processing_constraints_dict),
            )

    def _add_scheduler_metrics(
        self,
        benchmark: GenerativeBenchmark,
        headers: list[list[str]],
        values: list[str | int | float],
    ) -> None:
        """
        Add scheduler performance metrics including delays and processing times.

        :param benchmark: Benchmark data to extract scheduler metrics from
        :param headers: List of header hierarchies to append to
        :param values: List of values to append to
        """
        metrics = benchmark.scheduler_metrics

        requests_made_fields: list[tuple[str, int]] = [
            ("Requests Made Successful", metrics.requests_made.successful),
            ("Requests Made Incomplete", metrics.requests_made.incomplete),
            ("Requests Made Errored", metrics.requests_made.errored),
            ("Requests Made Total", metrics.requests_made.total),
        ]
        for field_name, value in requests_made_fields:
            self._add_field(headers, values, "Scheduler Metrics", field_name, value)

        timing_metrics: list[tuple[str, float]] = [
            ("Queued Time Avg", metrics.queued_time_avg),
            ("Resolve Start Delay Avg", metrics.resolve_start_delay_avg),
            (
                "Resolve Targeted Start Delay Avg",
                metrics.resolve_targeted_start_delay_avg,
            ),
            ("Request Start Delay Avg", metrics.request_start_delay_avg),
            (
                "Request Targeted Start Delay Avg",
                metrics.request_targeted_start_delay_avg,
            ),
            ("Request Time Avg", metrics.request_time_avg),
            ("Resolve End Delay Avg", metrics.resolve_end_delay_avg),
            ("Resolve Time Avg", metrics.resolve_time_avg),
            ("Finalized Delay Avg", metrics.finalized_delay_avg),
            ("Processed Delay Avg", metrics.processed_delay_avg),
        ]
        for field_name, timing in timing_metrics:
            self._add_field(
                headers, values, "Scheduler Metrics", field_name, timing, "Sec"
            )

        self._add_stats_for_metric(
            headers,
            values,
            metrics.generation_delay,
            "Generation Delay",
            "Sec",
        )

    def _add_stats_for_metric(
        self,
        headers: list[list[str]],
        values: list[str | int | float],
        metric: StatusDistributionSummary | DistributionSummary,
        group: str,
        units: str,
    ) -> None:
        """
        Add statistical summaries for a metric across all statuses.

        :param headers: List of header hierarchies to append to
        :param values: List of values to append to
        :param metric: Distribution summary to extract statistics from
        :param group: Top-level category for the metric
        :param units: Units for the metric values
        """
        if isinstance(metric, StatusDistributionSummary):
            for status in ["successful", "incomplete", "errored"]:
                dist = getattr(metric, status, None)
                if dist is None or dist.total_sum == 0.0:
                    continue
                self._add_distribution_stats(
                    headers, values, dist, group, units, status
                )
        else:
            self._add_distribution_stats(headers, values, metric, group, units, None)

    def _add_distribution_stats(
        self,
        headers: list[list[str]],
        values: list[str | int | float],
        dist: DistributionSummary,
        group: str,
        units: str,
        status: str | None,
    ) -> None:
        """
        Add distribution statistics including mean, median, and percentiles.

        :param headers: List of header hierarchies to append to
        :param values: List of values to append to
        :param dist: Distribution summary with statistical data
        :param group: Top-level category for the metric
        :param units: Units for the metric values
        :param status: Request status (successful, incomplete, errored) or None
        """
        status_prefix = f"{status.capitalize()} " if status else ""

        headers.append([group, f"{status_prefix}{units}", "Mean"])
        values.append(dist.mean)

        headers.append([group, f"{status_prefix}{units}", "Median"])
        values.append(dist.median)

        headers.append([group, f"{status_prefix}{units}", "Std Dev"])
        values.append(dist.std_dev)

        headers.append([group, f"{status_prefix}{units}", "Percentiles"])
        percentiles_str = (
            f"[{dist.min}, {dist.percentiles.p001}, {dist.percentiles.p01}, "
            f"{dist.percentiles.p05}, {dist.percentiles.p10}, {dist.percentiles.p25}, "
            f"{dist.percentiles.p75}, {dist.percentiles.p90}, {dist.percentiles.p95}, "
            f"{dist.percentiles.p99}, {dist.max}]"
        )
        values.append(percentiles_str)

finalize(report) async

Save the benchmark report as a CSV file.

Parameters:

Name Type Description Default
report GenerativeBenchmarksReport

The completed benchmark report

required

Returns:

Type Description
Path

Path to the saved CSV file

Source code in src/guidellm/benchmark/outputs/csv.py
async def finalize(self, report: GenerativeBenchmarksReport) -> Path:
    """
    Save the benchmark report as a CSV file.

    :param report: The completed benchmark report
    :return: Path to the saved CSV file
    """
    output_path = self.output_path
    if output_path.is_dir():
        output_path = output_path / GenerativeBenchmarkerCSV.DEFAULT_FILE
    output_path.parent.mkdir(parents=True, exist_ok=True)

    with output_path.open("w", newline="") as file:
        writer = csv.writer(file)

        all_headers: list[list[list[str]]] = []
        all_values: list[list[str | int | float]] = []

        for benchmark in report.benchmarks:
            benchmark_headers: list[list[str]] = []
            benchmark_values: list[str | int | float] = []

            self._add_run_info(benchmark, benchmark_headers, benchmark_values)
            self._add_benchmark_info(benchmark, benchmark_headers, benchmark_values)
            self._add_timing_info(benchmark, benchmark_headers, benchmark_values)
            self._add_request_counts(benchmark, benchmark_headers, benchmark_values)
            self._add_request_latency_metrics(
                benchmark, benchmark_headers, benchmark_values
            )
            self._add_server_throughput_metrics(
                benchmark, benchmark_headers, benchmark_values
            )
            for modality_name in MODALITY_METRICS:
                self._add_modality_metrics(
                    benchmark,
                    modality_name,
                    benchmark_headers,
                    benchmark_values,
                )
            self._add_scheduler_info(benchmark, benchmark_headers, benchmark_values)
            self._add_runtime_info(report, benchmark_headers, benchmark_values)
            self._add_interval_columns(
                benchmark, benchmark_headers, benchmark_values
            )

            all_headers.append(benchmark_headers)
            all_values.append(benchmark_values)

        headers, data_rows = self._align_columns(all_headers, all_values)

        self._write_multirow_header(writer, headers)
        for row in data_rows:
            writer.writerow(row)

    return output_path

from_args(args) classmethod

Create a CSV output formatter from output arguments.

Parameters:

Name Type Description Default
args BenchmarkOutputArgs

Output configuration with path

required

Returns:

Type Description
GenerativeBenchmarkerCSV

Configured CSV output formatter

Source code in src/guidellm/benchmark/outputs/csv.py
@classmethod
def from_args(cls, args: BenchmarkOutputArgs) -> GenerativeBenchmarkerCSV:
    """
    Create a CSV output formatter from output arguments.

    :param args: Output configuration with path
    :return: Configured CSV output formatter
    """
    if not isinstance(args, CSVBenchmarkOutputArgs):
        raise ValueError(f"Expected CSVBenchmarkOutputArgs, got {type(args)}")

    return cls(output_path=args.path)

GenerativeBenchmarkerConsole

Bases: GenerativeBenchmarkerOutput

Console output formatter for benchmark reports.

Renders benchmark results as formatted tables in the terminal, organizing metrics by category (run summary, request counts, latency, throughput, modality-specific) with proper alignment and type-specific formatting for readability.

Source code in src/guidellm/benchmark/outputs/console.py
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
631
632
633
634
635
636
637
638
639
640
641
642
643
644
645
646
647
648
649
650
651
652
653
654
655
656
657
658
659
660
661
662
663
664
665
666
667
668
669
670
671
672
673
674
675
676
677
678
679
680
681
682
683
684
685
686
687
688
689
690
691
692
693
694
695
696
697
698
699
700
701
702
703
704
705
706
707
708
709
710
711
712
713
714
715
716
717
718
719
720
721
722
723
724
725
726
727
728
729
730
731
732
733
734
735
736
737
738
739
740
741
742
743
744
745
746
747
748
749
750
751
752
753
754
755
756
757
758
759
760
761
762
763
764
765
766
767
768
769
770
771
772
773
774
775
776
777
778
779
780
781
782
783
784
785
786
787
788
789
790
791
792
793
794
795
796
797
798
799
800
801
802
803
804
805
806
807
@GenerativeBenchmarkerOutput.register("console")
class GenerativeBenchmarkerConsole(GenerativeBenchmarkerOutput):
    """
    Console output formatter for benchmark reports.

    Renders benchmark results as formatted tables in the terminal, organizing metrics
    by category (run summary, request counts, latency, throughput, modality-specific)
    with proper alignment and type-specific formatting for readability.
    """

    @classmethod
    def from_args(cls, _args: BenchmarkOutputArgs) -> GenerativeBenchmarkerConsole:
        """
        Create a console output formatter from output arguments.

        :param _args: Output configuration (unused for console output)
        :return: Configured console output formatter
        """
        return cls()

    console: Console = Field(
        default_factory=Console,
        description="Console utility for rendering formatted tables",
    )

    async def finalize(self, report: GenerativeBenchmarksReport) -> str:
        """
        Print the complete benchmark report to the console.

        Renders all metric tables including run summary, request counts, latency,
        throughput, and modality-specific statistics to the console.

        :param report: The completed benchmark report
        :return: Status message indicating output location
        """
        self.print_run_summary_table(report)
        self.print_text_table(report)
        self.print_image_table(report)
        self.print_video_table(report)
        self.print_audio_table(report)
        self.print_tool_call_table(report)
        self.print_request_counts_table(report)
        self.print_request_latency_table(report)
        self.print_server_throughput_table(report)

        return "printed to console"

    def print_run_summary_table(self, report: GenerativeBenchmarksReport):
        """
        Print the run summary table with timing and token information.

        :param report: The benchmark report containing run metadata
        """
        columns = ConsoleTableColumnsCollection()

        for benchmark in report.benchmarks:
            columns.add_value(
                benchmark.config.strategy.type_,
                group="Benchmark",
                name="Strategy",
                type_="text",
            )
            columns.add_value(
                benchmark.start_time, group="Timings", name="Start", type_="timestamp"
            )
            columns.add_value(
                benchmark.end_time, group="Timings", name="End", type_="timestamp"
            )
            columns.add_value(
                benchmark.duration, group="Timings", name="Dur", units="Sec"
            )
            columns.add_value(
                benchmark.warmup_duration, group="Timings", name="Warm", units="Sec"
            )
            columns.add_value(
                benchmark.cooldown_duration, group="Timings", name="Cool", units="Sec"
            )

            request_totals = benchmark.metrics.request_totals
            for count, name in (
                (request_totals.successful, "Comp"),
                (request_totals.incomplete, "Inc"),
                (request_totals.errored, "Err"),
            ):
                columns.add_value(
                    count,
                    group="Requests",
                    name=name,
                    units="Tot",
                    precision=0,
                )

            for token_metrics, group in [
                (benchmark.metrics.prompt_token_count, "Input Tokens"),
                (benchmark.metrics.output_token_count, "Output Tokens"),
            ]:
                columns.add_value(
                    token_metrics.successful.total_sum,
                    group=group,
                    name="Comp",
                    units="Tot",
                )
                columns.add_value(
                    token_metrics.incomplete.total_sum,
                    group=group,
                    name="Inc",
                    units="Tot",
                )
                columns.add_value(
                    token_metrics.errored.total_sum,
                    group=group,
                    name="Err",
                    units="Tot",
                )

        headers, values = columns.get_table_data()
        self.console.print("\n")
        self.console.print_table(headers, values, title="Run Summary Info")

    def print_text_table(self, report: GenerativeBenchmarksReport):
        """
        Print text-specific metrics table if any text data exists.

        :param report: The benchmark report containing text metrics
        """
        self._print_modality_table(
            report=report,
            modality="text",
            title="Text Metrics Statistics (Completed Requests)",
            metric_groups=[
                ("tokens", "Tokens"),
                ("words", "Words"),
                ("characters", "Characters"),
            ],
        )

    def print_image_table(self, report: GenerativeBenchmarksReport):
        """
        Print image-specific metrics table if any image data exists.

        :param report: The benchmark report containing image metrics
        """
        self._print_modality_table(
            report=report,
            modality="image",
            title="Image Metrics Statistics (Completed Requests)",
            metric_groups=[
                ("tokens", "Tokens"),
                ("images", "Images"),
                ("pixels", "Pixels"),
                ("bytes", "Bytes"),
            ],
        )

    def print_video_table(self, report: GenerativeBenchmarksReport):
        """
        Print video-specific metrics table if any video data exists.

        :param report: The benchmark report containing video metrics
        """
        self._print_modality_table(
            report=report,
            modality="video",
            title="Video Metrics Statistics (Completed Requests)",
            metric_groups=[
                ("tokens", "Tokens"),
                ("frames", "Frames"),
                ("seconds", "Seconds"),
                ("bytes", "Bytes"),
            ],
        )

    def print_audio_table(self, report: GenerativeBenchmarksReport):
        """
        Print audio-specific metrics table if any audio data exists.

        :param report: The benchmark report containing audio metrics
        """
        self._print_modality_table(
            report=report,
            modality="audio",
            title="Audio Metrics Statistics (Completed Requests)",
            metric_groups=[
                ("tokens", "Tokens"),
                ("samples", "Samples"),
                ("seconds", "Seconds"),
                ("bytes", "Bytes"),
            ],
        )

    def print_tool_call_table(self, report: GenerativeBenchmarksReport):
        """
        Print tool-call-specific metrics table if any tool call data exists.

        :param report: The benchmark report containing tool call metrics
        """
        self._print_modality_table(
            report=report,
            modality="tool_call",
            title="Tool Call Metrics Statistics (Completed Requests)",
            metric_groups=[
                ("tokens", "Tokens"),
                ("mixed_tokens", "Mixed Tokens"),
                ("count", "Count"),
            ],
        )

    def print_request_counts_table(self, report: GenerativeBenchmarksReport):
        """
        Print request token count statistics table.

        :param report: The benchmark report containing request count metrics
        """
        columns = ConsoleTableColumnsCollection()

        for benchmark in report.benchmarks:
            columns.add_value(
                benchmark.config.strategy.type_,
                group="Benchmark",
                name="Strategy",
                type_="text",
            )
            columns.add_stats(
                benchmark.metrics.prompt_token_count,
                group="Input Tok",
                name="Per Req",
            )
            columns.add_stats(
                benchmark.metrics.output_token_count,
                group="Output Tok",
                name="Per Req",
            )
            columns.add_stats(
                benchmark.metrics.total_token_count,
                group="Total Tok",
                name="Per Req",
            )
            columns.add_stats(
                benchmark.metrics.request_streaming_iterations_count,
                group="Stream Iter",
                name="Per Req",
            )
            columns.add_stats(
                benchmark.metrics.output_tokens_per_iteration,
                group="Output Tok",
                name="Per Stream Iter",
            )

        headers, values = columns.get_table_data()
        self.console.print("\n")
        self.console.print_table(
            headers,
            values,
            title="Request Token Statistics (Completed Requests)",
        )

    def print_request_latency_table(self, report: GenerativeBenchmarksReport):
        """
        Print request latency metrics table.

        :param report: The benchmark report containing latency metrics
        """
        columns = ConsoleTableColumnsCollection()

        for benchmark in report.benchmarks:
            columns.add_value(
                benchmark.config.strategy.type_,
                group="Benchmark",
                name="Strategy",
                type_="text",
            )
            # ITL and TPOT are reported without a margin because their mean is
            # a ratio over output tokens rather than a mean over requests; the
            # column shows the value alone for them.
            latency_stats: tuple[StatTypesAlias, ...] = (
                "mean_moe",
                "median",
                "p95_ci",
            )
            columns.add_stats(
                benchmark.metrics.request_latency,
                group="Request Latency",
                name="Sec",
                types=latency_stats,
            )
            columns.add_stats(
                benchmark.metrics.time_to_first_token_ms,
                group="TTFT",
                name="ms",
                types=latency_stats,
            )
            columns.add_stats(
                benchmark.metrics.time_to_first_output_token_ms,
                group="TTFOT",
                name="ms",
                types=latency_stats,
            )
            columns.add_stats(
                benchmark.metrics.inter_token_latency_ms,
                group="ITL",
                name="ms",
                types=latency_stats,
            )
            columns.add_stats(
                benchmark.metrics.time_per_output_token_ms,
                group="TPOT",
                name="ms",
                types=latency_stats,
            )
        headers, values = columns.get_table_data()
        self.console.print("\n")
        self.console.print_table(
            headers,
            values,
            title="Request Latency Statistics (Completed Requests)",
        )
        if any(
            isinstance(value, str) and value.endswith(UNSUPPORTED_PERCENTILE_MARKER)
            for column in values
            for value in column
        ):
            self.console.print(UNSUPPORTED_PERCENTILE_FOOTNOTE)

    def print_server_throughput_table(self, report: GenerativeBenchmarksReport):
        """
        Print server throughput metrics table.

        :param report: The benchmark report containing throughput metrics
        """
        columns = ConsoleTableColumnsCollection()
        # Widen the table whenever objectives were configured, not only when
        # they could be evaluated. A workload that cannot measure an objective,
        # such as time to first token without streaming, then shows empty cells
        # instead of the table silently looking as though none were set.
        report_has_goodput = any(
            benchmark.config.slo is not None for benchmark in report.benchmarks
        )

        for benchmark in report.benchmarks:
            columns.add_value(
                benchmark.config.strategy.type_,
                group="Benchmark",
                name="Strategy",
                type_="text",
            )
            columns.add_stats(
                benchmark.metrics.request_concurrency,
                status="total",
                group="Requests",
                name="Concurrency",
                types=("median", "mean"),
            )
            columns.add_stats(
                benchmark.metrics.requests_per_second,
                status="total",
                group="Requests",
                name="Per Sec",
                types=("mean",),
            )
            columns.add_stats(
                benchmark.metrics.prompt_tokens_per_second,
                status="total",
                group="Input Tokens",
                name="Per Sec",
                types=("mean",),
            )
            columns.add_stats(
                benchmark.metrics.output_tokens_per_second,
                status="total",
                group="Output Tokens",
                name="Per Sec",
                types=("mean",),
            )
            columns.add_stats(
                benchmark.metrics.tokens_per_second,
                status="total",
                group="Total Tokens",
                name="Per Sec",
                types=("mean",),
            )
            if report_has_goodput:
                attainment = benchmark.metrics.slo_attainment
                columns.add_value(
                    None if attainment is None else attainment * 100.0,
                    group="Goodput",
                    name="Attainment",
                    units="%",
                    precision=1,
                )
                columns.add_stats(
                    benchmark.metrics.request_goodput,
                    status="total",
                    group="Goodput",
                    name="Per Sec",
                    types=("mean",),
                )

        headers, values = columns.get_table_data()
        self.console.print("\n")
        self.console.print_table(
            headers, values, title="Server Throughput Statistics (All Requests)"
        )

    def _print_modality_table(
        self,
        report: GenerativeBenchmarksReport,
        modality: Literal["text", "image", "video", "audio", "tool_call"],
        title: str,
        metric_groups: list[tuple[str, str]],
    ):
        columns: dict[str, ConsoleTableColumnsCollection] = defaultdict(
            ConsoleTableColumnsCollection
        )

        for benchmark in report.benchmarks:
            columns["labels"].add_value(
                benchmark.config.strategy.type_,
                group="Benchmark",
                name="Strategy",
                type_="text",
            )

            modality_metrics = getattr(benchmark.metrics, modality)

            for metric_attr, display_name in metric_groups:
                metric_obj = getattr(modality_metrics, metric_attr, None)
                input_stats: StatusDistributionSummary | None = (
                    getattr(metric_obj, "input", None) if metric_obj else None
                )
                columns[f"{metric_attr}.input"].add_stats(
                    input_stats,
                    group=f"Input {display_name}",
                    name="Per Request",
                )
                input_per_second_stats: StatusDistributionSummary | None = (
                    getattr(metric_obj, "input_per_second", None)
                    if metric_obj
                    else None
                )
                columns[f"{metric_attr}.input"].add_stats(
                    input_per_second_stats,
                    group=f"Input {display_name}",
                    name="Per Second",
                    types=("median", "mean"),
                )
                output_stats: StatusDistributionSummary | None = (
                    getattr(metric_obj, "output", None) if metric_obj else None
                )
                columns[f"{metric_attr}.output"].add_stats(
                    output_stats,
                    group=f"Output {display_name}",
                    name="Per Request",
                )
                output_per_second_stats: StatusDistributionSummary | None = (
                    getattr(metric_obj, "output_per_second", None)
                    if metric_obj
                    else None
                )
                columns[f"{metric_attr}.output"].add_stats(
                    output_per_second_stats,
                    group=f"Output {display_name}",
                    name="Per Second",
                    types=("median", "mean"),
                )

        self._print_inp_out_tables(
            title=title,
            labels=columns["labels"],
            groups=[
                (columns[f"{metric_attr}.input"], columns[f"{metric_attr}.output"])
                for metric_attr, _ in metric_groups
            ],
        )

    def _print_inp_out_tables(
        self,
        title: str,
        labels: ConsoleTableColumnsCollection,
        groups: list[
            tuple[ConsoleTableColumnsCollection, ConsoleTableColumnsCollection]
        ],
    ):
        input_headers, input_values = [], []
        output_headers, output_values = [], []
        input_has_data = False
        output_has_data = False

        for input_columns, output_columns in groups:
            # Check if columns have any non-None values
            type_input_has_data = any(
                any(value is not None for value in column.values)
                for column in input_columns.values()
            )
            type_output_has_data = any(
                any(value is not None for value in column.values)
                for column in output_columns.values()
            )

            if not (type_input_has_data or type_output_has_data):
                continue

            input_has_data = input_has_data or type_input_has_data
            output_has_data = output_has_data or type_output_has_data

            input_type_headers, input_type_columns = input_columns.get_table_data()
            output_type_headers, output_type_columns = output_columns.get_table_data()

            input_headers.extend(input_type_headers)
            input_values.extend(input_type_columns)
            output_headers.extend(output_type_headers)
            output_values.extend(output_type_columns)

        if not (input_has_data or output_has_data):
            return

        labels_headers, labels_values = labels.get_table_data()
        header_cols_groups = []
        value_cols_groups = []

        if input_has_data:
            header_cols_groups.append(labels_headers + input_headers)
            value_cols_groups.append(labels_values + input_values)
        if output_has_data:
            header_cols_groups.append(labels_headers + output_headers)
            value_cols_groups.append(labels_values + output_values)

        if header_cols_groups and value_cols_groups:
            self.console.print("\n")
            self.console.print_tables(
                header_cols_groups=header_cols_groups,
                value_cols_groups=value_cols_groups,
                title=title,
            )

finalize(report) async

Print the complete benchmark report to the console.

Renders all metric tables including run summary, request counts, latency, throughput, and modality-specific statistics to the console.

Parameters:

Name Type Description Default
report GenerativeBenchmarksReport

The completed benchmark report

required

Returns:

Type Description
str

Status message indicating output location

Source code in src/guidellm/benchmark/outputs/console.py
async def finalize(self, report: GenerativeBenchmarksReport) -> str:
    """
    Print the complete benchmark report to the console.

    Renders all metric tables including run summary, request counts, latency,
    throughput, and modality-specific statistics to the console.

    :param report: The completed benchmark report
    :return: Status message indicating output location
    """
    self.print_run_summary_table(report)
    self.print_text_table(report)
    self.print_image_table(report)
    self.print_video_table(report)
    self.print_audio_table(report)
    self.print_tool_call_table(report)
    self.print_request_counts_table(report)
    self.print_request_latency_table(report)
    self.print_server_throughput_table(report)

    return "printed to console"

from_args(_args) classmethod

Create a console output formatter from output arguments.

Parameters:

Name Type Description Default
_args BenchmarkOutputArgs

Output configuration (unused for console output)

required

Returns:

Type Description
GenerativeBenchmarkerConsole

Configured console output formatter

Source code in src/guidellm/benchmark/outputs/console.py
@classmethod
def from_args(cls, _args: BenchmarkOutputArgs) -> GenerativeBenchmarkerConsole:
    """
    Create a console output formatter from output arguments.

    :param _args: Output configuration (unused for console output)
    :return: Configured console output formatter
    """
    return cls()

print_audio_table(report)

Print audio-specific metrics table if any audio data exists.

Parameters:

Name Type Description Default
report GenerativeBenchmarksReport

The benchmark report containing audio metrics

required
Source code in src/guidellm/benchmark/outputs/console.py
def print_audio_table(self, report: GenerativeBenchmarksReport):
    """
    Print audio-specific metrics table if any audio data exists.

    :param report: The benchmark report containing audio metrics
    """
    self._print_modality_table(
        report=report,
        modality="audio",
        title="Audio Metrics Statistics (Completed Requests)",
        metric_groups=[
            ("tokens", "Tokens"),
            ("samples", "Samples"),
            ("seconds", "Seconds"),
            ("bytes", "Bytes"),
        ],
    )

print_image_table(report)

Print image-specific metrics table if any image data exists.

Parameters:

Name Type Description Default
report GenerativeBenchmarksReport

The benchmark report containing image metrics

required
Source code in src/guidellm/benchmark/outputs/console.py
def print_image_table(self, report: GenerativeBenchmarksReport):
    """
    Print image-specific metrics table if any image data exists.

    :param report: The benchmark report containing image metrics
    """
    self._print_modality_table(
        report=report,
        modality="image",
        title="Image Metrics Statistics (Completed Requests)",
        metric_groups=[
            ("tokens", "Tokens"),
            ("images", "Images"),
            ("pixels", "Pixels"),
            ("bytes", "Bytes"),
        ],
    )

print_request_counts_table(report)

Print request token count statistics table.

Parameters:

Name Type Description Default
report GenerativeBenchmarksReport

The benchmark report containing request count metrics

required
Source code in src/guidellm/benchmark/outputs/console.py
def print_request_counts_table(self, report: GenerativeBenchmarksReport):
    """
    Print request token count statistics table.

    :param report: The benchmark report containing request count metrics
    """
    columns = ConsoleTableColumnsCollection()

    for benchmark in report.benchmarks:
        columns.add_value(
            benchmark.config.strategy.type_,
            group="Benchmark",
            name="Strategy",
            type_="text",
        )
        columns.add_stats(
            benchmark.metrics.prompt_token_count,
            group="Input Tok",
            name="Per Req",
        )
        columns.add_stats(
            benchmark.metrics.output_token_count,
            group="Output Tok",
            name="Per Req",
        )
        columns.add_stats(
            benchmark.metrics.total_token_count,
            group="Total Tok",
            name="Per Req",
        )
        columns.add_stats(
            benchmark.metrics.request_streaming_iterations_count,
            group="Stream Iter",
            name="Per Req",
        )
        columns.add_stats(
            benchmark.metrics.output_tokens_per_iteration,
            group="Output Tok",
            name="Per Stream Iter",
        )

    headers, values = columns.get_table_data()
    self.console.print("\n")
    self.console.print_table(
        headers,
        values,
        title="Request Token Statistics (Completed Requests)",
    )

print_request_latency_table(report)

Print request latency metrics table.

Parameters:

Name Type Description Default
report GenerativeBenchmarksReport

The benchmark report containing latency metrics

required
Source code in src/guidellm/benchmark/outputs/console.py
def print_request_latency_table(self, report: GenerativeBenchmarksReport):
    """
    Print request latency metrics table.

    :param report: The benchmark report containing latency metrics
    """
    columns = ConsoleTableColumnsCollection()

    for benchmark in report.benchmarks:
        columns.add_value(
            benchmark.config.strategy.type_,
            group="Benchmark",
            name="Strategy",
            type_="text",
        )
        # ITL and TPOT are reported without a margin because their mean is
        # a ratio over output tokens rather than a mean over requests; the
        # column shows the value alone for them.
        latency_stats: tuple[StatTypesAlias, ...] = (
            "mean_moe",
            "median",
            "p95_ci",
        )
        columns.add_stats(
            benchmark.metrics.request_latency,
            group="Request Latency",
            name="Sec",
            types=latency_stats,
        )
        columns.add_stats(
            benchmark.metrics.time_to_first_token_ms,
            group="TTFT",
            name="ms",
            types=latency_stats,
        )
        columns.add_stats(
            benchmark.metrics.time_to_first_output_token_ms,
            group="TTFOT",
            name="ms",
            types=latency_stats,
        )
        columns.add_stats(
            benchmark.metrics.inter_token_latency_ms,
            group="ITL",
            name="ms",
            types=latency_stats,
        )
        columns.add_stats(
            benchmark.metrics.time_per_output_token_ms,
            group="TPOT",
            name="ms",
            types=latency_stats,
        )
    headers, values = columns.get_table_data()
    self.console.print("\n")
    self.console.print_table(
        headers,
        values,
        title="Request Latency Statistics (Completed Requests)",
    )
    if any(
        isinstance(value, str) and value.endswith(UNSUPPORTED_PERCENTILE_MARKER)
        for column in values
        for value in column
    ):
        self.console.print(UNSUPPORTED_PERCENTILE_FOOTNOTE)

print_run_summary_table(report)

Print the run summary table with timing and token information.

Parameters:

Name Type Description Default
report GenerativeBenchmarksReport

The benchmark report containing run metadata

required
Source code in src/guidellm/benchmark/outputs/console.py
def print_run_summary_table(self, report: GenerativeBenchmarksReport):
    """
    Print the run summary table with timing and token information.

    :param report: The benchmark report containing run metadata
    """
    columns = ConsoleTableColumnsCollection()

    for benchmark in report.benchmarks:
        columns.add_value(
            benchmark.config.strategy.type_,
            group="Benchmark",
            name="Strategy",
            type_="text",
        )
        columns.add_value(
            benchmark.start_time, group="Timings", name="Start", type_="timestamp"
        )
        columns.add_value(
            benchmark.end_time, group="Timings", name="End", type_="timestamp"
        )
        columns.add_value(
            benchmark.duration, group="Timings", name="Dur", units="Sec"
        )
        columns.add_value(
            benchmark.warmup_duration, group="Timings", name="Warm", units="Sec"
        )
        columns.add_value(
            benchmark.cooldown_duration, group="Timings", name="Cool", units="Sec"
        )

        request_totals = benchmark.metrics.request_totals
        for count, name in (
            (request_totals.successful, "Comp"),
            (request_totals.incomplete, "Inc"),
            (request_totals.errored, "Err"),
        ):
            columns.add_value(
                count,
                group="Requests",
                name=name,
                units="Tot",
                precision=0,
            )

        for token_metrics, group in [
            (benchmark.metrics.prompt_token_count, "Input Tokens"),
            (benchmark.metrics.output_token_count, "Output Tokens"),
        ]:
            columns.add_value(
                token_metrics.successful.total_sum,
                group=group,
                name="Comp",
                units="Tot",
            )
            columns.add_value(
                token_metrics.incomplete.total_sum,
                group=group,
                name="Inc",
                units="Tot",
            )
            columns.add_value(
                token_metrics.errored.total_sum,
                group=group,
                name="Err",
                units="Tot",
            )

    headers, values = columns.get_table_data()
    self.console.print("\n")
    self.console.print_table(headers, values, title="Run Summary Info")

print_server_throughput_table(report)

Print server throughput metrics table.

Parameters:

Name Type Description Default
report GenerativeBenchmarksReport

The benchmark report containing throughput metrics

required
Source code in src/guidellm/benchmark/outputs/console.py
def print_server_throughput_table(self, report: GenerativeBenchmarksReport):
    """
    Print server throughput metrics table.

    :param report: The benchmark report containing throughput metrics
    """
    columns = ConsoleTableColumnsCollection()
    # Widen the table whenever objectives were configured, not only when
    # they could be evaluated. A workload that cannot measure an objective,
    # such as time to first token without streaming, then shows empty cells
    # instead of the table silently looking as though none were set.
    report_has_goodput = any(
        benchmark.config.slo is not None for benchmark in report.benchmarks
    )

    for benchmark in report.benchmarks:
        columns.add_value(
            benchmark.config.strategy.type_,
            group="Benchmark",
            name="Strategy",
            type_="text",
        )
        columns.add_stats(
            benchmark.metrics.request_concurrency,
            status="total",
            group="Requests",
            name="Concurrency",
            types=("median", "mean"),
        )
        columns.add_stats(
            benchmark.metrics.requests_per_second,
            status="total",
            group="Requests",
            name="Per Sec",
            types=("mean",),
        )
        columns.add_stats(
            benchmark.metrics.prompt_tokens_per_second,
            status="total",
            group="Input Tokens",
            name="Per Sec",
            types=("mean",),
        )
        columns.add_stats(
            benchmark.metrics.output_tokens_per_second,
            status="total",
            group="Output Tokens",
            name="Per Sec",
            types=("mean",),
        )
        columns.add_stats(
            benchmark.metrics.tokens_per_second,
            status="total",
            group="Total Tokens",
            name="Per Sec",
            types=("mean",),
        )
        if report_has_goodput:
            attainment = benchmark.metrics.slo_attainment
            columns.add_value(
                None if attainment is None else attainment * 100.0,
                group="Goodput",
                name="Attainment",
                units="%",
                precision=1,
            )
            columns.add_stats(
                benchmark.metrics.request_goodput,
                status="total",
                group="Goodput",
                name="Per Sec",
                types=("mean",),
            )

    headers, values = columns.get_table_data()
    self.console.print("\n")
    self.console.print_table(
        headers, values, title="Server Throughput Statistics (All Requests)"
    )

print_text_table(report)

Print text-specific metrics table if any text data exists.

Parameters:

Name Type Description Default
report GenerativeBenchmarksReport

The benchmark report containing text metrics

required
Source code in src/guidellm/benchmark/outputs/console.py
def print_text_table(self, report: GenerativeBenchmarksReport):
    """
    Print text-specific metrics table if any text data exists.

    :param report: The benchmark report containing text metrics
    """
    self._print_modality_table(
        report=report,
        modality="text",
        title="Text Metrics Statistics (Completed Requests)",
        metric_groups=[
            ("tokens", "Tokens"),
            ("words", "Words"),
            ("characters", "Characters"),
        ],
    )

print_tool_call_table(report)

Print tool-call-specific metrics table if any tool call data exists.

Parameters:

Name Type Description Default
report GenerativeBenchmarksReport

The benchmark report containing tool call metrics

required
Source code in src/guidellm/benchmark/outputs/console.py
def print_tool_call_table(self, report: GenerativeBenchmarksReport):
    """
    Print tool-call-specific metrics table if any tool call data exists.

    :param report: The benchmark report containing tool call metrics
    """
    self._print_modality_table(
        report=report,
        modality="tool_call",
        title="Tool Call Metrics Statistics (Completed Requests)",
        metric_groups=[
            ("tokens", "Tokens"),
            ("mixed_tokens", "Mixed Tokens"),
            ("count", "Count"),
        ],
    )

print_video_table(report)

Print video-specific metrics table if any video data exists.

Parameters:

Name Type Description Default
report GenerativeBenchmarksReport

The benchmark report containing video metrics

required
Source code in src/guidellm/benchmark/outputs/console.py
def print_video_table(self, report: GenerativeBenchmarksReport):
    """
    Print video-specific metrics table if any video data exists.

    :param report: The benchmark report containing video metrics
    """
    self._print_modality_table(
        report=report,
        modality="video",
        title="Video Metrics Statistics (Completed Requests)",
        metric_groups=[
            ("tokens", "Tokens"),
            ("frames", "Frames"),
            ("seconds", "Seconds"),
            ("bytes", "Bytes"),
        ],
    )

GenerativeBenchmarkerHTML

Bases: GenerativeBenchmarkerOutput

Self-contained HTML report formatter for generative benchmarks.

Embeds compact chart/table JSON into a packaged HTML/CSS/JS template so the resulting file can be shared without network access or versioned UI assets.

Attributes:

Name Type Description
DEFAULT_FILE str

Default filename when path is a directory

Source code in src/guidellm/benchmark/outputs/html.py
@GenerativeBenchmarkerOutput.register("html")
class GenerativeBenchmarkerHTML(GenerativeBenchmarkerOutput):
    """
    Self-contained HTML report formatter for generative benchmarks.

    Embeds compact chart/table JSON into a packaged HTML/CSS/JS template so the
    resulting file can be shared without network access or versioned UI assets.

    :cvar DEFAULT_FILE: Default filename when ``path`` is a directory
    """

    DEFAULT_FILE: ClassVar[str] = "benchmarks.html"

    output_path: Path = Field(
        default_factory=Path.cwd,
        description=(
            "Directory or file path for saving the HTML report, "
            "defaults to current working directory"
        ),
    )

    @classmethod
    def from_args(cls, args: BenchmarkOutputArgs) -> GenerativeBenchmarkerHTML:
        """
        Create an HTML output formatter from output arguments.

        :param args: Output configuration with path
        :return: Configured HTML output formatter
        """
        if not isinstance(args, HTMLBenchmarkOutputArgs):
            raise TypeError(f"Expected HTMLBenchmarkOutputArgs, got {type(args)}")

        return cls(output_path=args.path)

    async def finalize(self, report: GenerativeBenchmarksReport) -> Path:
        """
        Generate and save the self-contained HTML benchmark report.

        :param report: Completed benchmark report containing all results
        :return: Path to the saved HTML report file
        """
        output_path = self.output_path
        if output_path.is_dir():
            output_path = output_path / self.DEFAULT_FILE
        output_path.parent.mkdir(parents=True, exist_ok=True)

        view = build_report_view(report)
        html = render_html_report(view)
        output_path.write_text(html, encoding="utf-8")
        logger.debug("Saved HTML report to {}", output_path)
        return output_path

finalize(report) async

Generate and save the self-contained HTML benchmark report.

Parameters:

Name Type Description Default
report GenerativeBenchmarksReport

Completed benchmark report containing all results

required

Returns:

Type Description
Path

Path to the saved HTML report file

Source code in src/guidellm/benchmark/outputs/html.py
async def finalize(self, report: GenerativeBenchmarksReport) -> Path:
    """
    Generate and save the self-contained HTML benchmark report.

    :param report: Completed benchmark report containing all results
    :return: Path to the saved HTML report file
    """
    output_path = self.output_path
    if output_path.is_dir():
        output_path = output_path / self.DEFAULT_FILE
    output_path.parent.mkdir(parents=True, exist_ok=True)

    view = build_report_view(report)
    html = render_html_report(view)
    output_path.write_text(html, encoding="utf-8")
    logger.debug("Saved HTML report to {}", output_path)
    return output_path

from_args(args) classmethod

Create an HTML output formatter from output arguments.

Parameters:

Name Type Description Default
args BenchmarkOutputArgs

Output configuration with path

required

Returns:

Type Description
GenerativeBenchmarkerHTML

Configured HTML output formatter

Source code in src/guidellm/benchmark/outputs/html.py
@classmethod
def from_args(cls, args: BenchmarkOutputArgs) -> GenerativeBenchmarkerHTML:
    """
    Create an HTML output formatter from output arguments.

    :param args: Output configuration with path
    :return: Configured HTML output formatter
    """
    if not isinstance(args, HTMLBenchmarkOutputArgs):
        raise TypeError(f"Expected HTMLBenchmarkOutputArgs, got {type(args)}")

    return cls(output_path=args.path)

GenerativeBenchmarkerOutput

Bases: BaseModel, RegistryMixin[type['GenerativeBenchmarkerOutput']], ABC

Abstract base for benchmark output formatters with registry support.

Defines the interface for transforming benchmark reports into various output formats. Subclasses implement specific formatters (JSON, CSV, HTML) that can be registered and resolved dynamically.

Example: ::

    output = GenerativeBenchmarkerOutput.resolve(
        JSONBenchmarkOutputArgs(path="./results.json")
    )
    await output.finalize(report)
Source code in src/guidellm/benchmark/outputs/output.py
class GenerativeBenchmarkerOutput(
    BaseModel, RegistryMixin[type["GenerativeBenchmarkerOutput"]], ABC
):
    """
    Abstract base for benchmark output formatters with registry support.

    Defines the interface for transforming benchmark reports into various output
    formats. Subclasses implement specific formatters (JSON, CSV, HTML) that can be
    registered and resolved dynamically.

    Example:
        ::

            output = GenerativeBenchmarkerOutput.resolve(
                JSONBenchmarkOutputArgs(path="./results.json")
            )
            await output.finalize(report)
    """

    model_config = ConfigDict(
        extra="ignore",
        arbitrary_types_allowed=True,
        validate_assignment=True,
        from_attributes=True,
        use_enum_values=True,
    )

    @classmethod
    @abstractmethod
    def from_args(cls, args: BenchmarkOutputArgs) -> GenerativeBenchmarkerOutput:
        """
        Create an output formatter instance from output arguments.

        :param args: Output configuration arguments
        :return: Configured output formatter instance
        """
        ...

    @classmethod
    def resolve(cls, args: BenchmarkOutputArgs) -> GenerativeBenchmarkerOutput:
        """
        Resolve output arguments into a formatter instance.

        Looks up the registered output class by ``args.kind`` and delegates
        construction to its :meth:`from_args` factory.

        :param args: Output configuration arguments with kind and format-specific fields
        :return: Configured output formatter instance
        :raises ValueError: If the output kind is not registered
        """
        output_class = cls.get_registered_object(args.kind)
        if output_class is None:
            available_formats = list(cls.registry.keys()) if cls.registry else []
            raise ValueError(
                f"Output format '{args.kind}' is not registered. "
                f"Available formats: {available_formats}"
            )
        return output_class.from_args(args)

    @abstractmethod
    async def finalize(self, report: GenerativeBenchmarksReport) -> Any:
        """
        Process and persist benchmark report in the formatter's output format.

        :param report: Benchmark report containing results to format and output
        :return: Format-specific output result (file path, response object, etc.)
        """
        ...

finalize(report) abstractmethod async

Process and persist benchmark report in the formatter's output format.

Parameters:

Name Type Description Default
report GenerativeBenchmarksReport

Benchmark report containing results to format and output

required

Returns:

Type Description
Any

Format-specific output result (file path, response object, etc.)

Source code in src/guidellm/benchmark/outputs/output.py
@abstractmethod
async def finalize(self, report: GenerativeBenchmarksReport) -> Any:
    """
    Process and persist benchmark report in the formatter's output format.

    :param report: Benchmark report containing results to format and output
    :return: Format-specific output result (file path, response object, etc.)
    """
    ...

from_args(args) abstractmethod classmethod

Create an output formatter instance from output arguments.

Parameters:

Name Type Description Default
args BenchmarkOutputArgs

Output configuration arguments

required

Returns:

Type Description
GenerativeBenchmarkerOutput

Configured output formatter instance

Source code in src/guidellm/benchmark/outputs/output.py
@classmethod
@abstractmethod
def from_args(cls, args: BenchmarkOutputArgs) -> GenerativeBenchmarkerOutput:
    """
    Create an output formatter instance from output arguments.

    :param args: Output configuration arguments
    :return: Configured output formatter instance
    """
    ...

resolve(args) classmethod

Resolve output arguments into a formatter instance.

Looks up the registered output class by args.kind and delegates construction to its :meth:from_args factory.

Parameters:

Name Type Description Default
args BenchmarkOutputArgs

Output configuration arguments with kind and format-specific fields

required

Returns:

Type Description
GenerativeBenchmarkerOutput

Configured output formatter instance

Raises:

Type Description
ValueError

If the output kind is not registered

Source code in src/guidellm/benchmark/outputs/output.py
@classmethod
def resolve(cls, args: BenchmarkOutputArgs) -> GenerativeBenchmarkerOutput:
    """
    Resolve output arguments into a formatter instance.

    Looks up the registered output class by ``args.kind`` and delegates
    construction to its :meth:`from_args` factory.

    :param args: Output configuration arguments with kind and format-specific fields
    :return: Configured output formatter instance
    :raises ValueError: If the output kind is not registered
    """
    output_class = cls.get_registered_object(args.kind)
    if output_class is None:
        available_formats = list(cls.registry.keys()) if cls.registry else []
        raise ValueError(
            f"Output format '{args.kind}' is not registered. "
            f"Available formats: {available_formats}"
        )
    return output_class.from_args(args)

GenerativeBenchmarkerPlot

Bases: GenerativeBenchmarkerOutput

Plot output formatter for benchmark results.

Generates a high-quality dashboard visualization of LLM benchmark results, enforcing light theme and saving to a PNG image file.

Source code in src/guidellm/benchmark/outputs/plot.py
@GenerativeBenchmarkerOutput.register("plot")
class GenerativeBenchmarkerPlot(GenerativeBenchmarkerOutput):
    """
    Plot output formatter for benchmark results.

    Generates a high-quality dashboard visualization of LLM benchmark results,
    enforcing light theme and saving to a PNG image file.
    """

    output_path: Path = Field(
        default_factory=Path.cwd,
        description=(
            "Path where the PNG plot file will be saved, defaults to current directory"
        ),
    )
    dpi: int = Field(
        default=100,
        description="Resolution of the output image in Dots Per Inch.",
    )

    @classmethod
    def from_args(cls, args: BenchmarkOutputArgs) -> GenerativeBenchmarkerPlot:
        """
        Create a plot output formatter from output arguments.

        :param args: Output configuration with path and dpi
        :return: Configured plot output formatter
        """
        if not isinstance(args, PlotBenchmarkOutputArgs):
            raise ValueError(f"Expected PlotBenchmarkOutputArgs, got {type(args)}")

        return cls(output_path=args.path, dpi=args.dpi)

    async def finalize(self, report: GenerativeBenchmarksReport) -> Path:
        """
        Save the benchmark report as a static PNG plot visualization.

        :param report: The completed benchmark report
        :return: Path to the saved PNG plot file
        """
        output_path = self.output_path
        if output_path.is_dir():
            output_path = output_path / "benchmarks.png"

        output_path.parent.mkdir(parents=True, exist_ok=True)

        if not report.benchmarks:
            fig, ax = plot.plt.subplots(figsize=(10, 8), facecolor="white")
            ax.text(0.5, 0.5, "No benchmark data available", fontsize=14, ha="center")
            ax.set_facecolor("white")
            plot.plt.savefig(
                output_path, facecolor="white", bbox_inches="tight", dpi=self.dpi
            )
            plot.plt.close(fig)
            return output_path

        points = _build_points(report.benchmarks)

        _set_plot_theme(self.dpi)

        fig, axs = plot.plt.subplots(4, 2, figsize=(16, 22), facecolor="white")
        fig.suptitle(
            "GuideLLM Benchmark Performance Visualization",
            fontsize=18,
            color="black",
            weight="bold",
            y=0.985,
        )

        for ax in axs.flat:
            _style_axis(ax)

        for ax, plot_function in zip(axs.flat, _PLOT_FUNCTIONS, strict=True):
            plot_function(ax, points)

        fig.align_ylabels(axs[:, 0])
        fig.align_ylabels(axs[:, 1])
        plot.plt.tight_layout(rect=(0.04, 0.03, 0.96, 0.965), h_pad=3.0, w_pad=2.5)
        fig.savefig(output_path, facecolor="white", bbox_inches="tight", dpi=self.dpi)
        plot.plt.close(fig)

        return output_path

finalize(report) async

Save the benchmark report as a static PNG plot visualization.

Parameters:

Name Type Description Default
report GenerativeBenchmarksReport

The completed benchmark report

required

Returns:

Type Description
Path

Path to the saved PNG plot file

Source code in src/guidellm/benchmark/outputs/plot.py
async def finalize(self, report: GenerativeBenchmarksReport) -> Path:
    """
    Save the benchmark report as a static PNG plot visualization.

    :param report: The completed benchmark report
    :return: Path to the saved PNG plot file
    """
    output_path = self.output_path
    if output_path.is_dir():
        output_path = output_path / "benchmarks.png"

    output_path.parent.mkdir(parents=True, exist_ok=True)

    if not report.benchmarks:
        fig, ax = plot.plt.subplots(figsize=(10, 8), facecolor="white")
        ax.text(0.5, 0.5, "No benchmark data available", fontsize=14, ha="center")
        ax.set_facecolor("white")
        plot.plt.savefig(
            output_path, facecolor="white", bbox_inches="tight", dpi=self.dpi
        )
        plot.plt.close(fig)
        return output_path

    points = _build_points(report.benchmarks)

    _set_plot_theme(self.dpi)

    fig, axs = plot.plt.subplots(4, 2, figsize=(16, 22), facecolor="white")
    fig.suptitle(
        "GuideLLM Benchmark Performance Visualization",
        fontsize=18,
        color="black",
        weight="bold",
        y=0.985,
    )

    for ax in axs.flat:
        _style_axis(ax)

    for ax, plot_function in zip(axs.flat, _PLOT_FUNCTIONS, strict=True):
        plot_function(ax, points)

    fig.align_ylabels(axs[:, 0])
    fig.align_ylabels(axs[:, 1])
    plot.plt.tight_layout(rect=(0.04, 0.03, 0.96, 0.965), h_pad=3.0, w_pad=2.5)
    fig.savefig(output_path, facecolor="white", bbox_inches="tight", dpi=self.dpi)
    plot.plt.close(fig)

    return output_path

from_args(args) classmethod

Create a plot output formatter from output arguments.

Parameters:

Name Type Description Default
args BenchmarkOutputArgs

Output configuration with path and dpi

required

Returns:

Type Description
GenerativeBenchmarkerPlot

Configured plot output formatter

Source code in src/guidellm/benchmark/outputs/plot.py
@classmethod
def from_args(cls, args: BenchmarkOutputArgs) -> GenerativeBenchmarkerPlot:
    """
    Create a plot output formatter from output arguments.

    :param args: Output configuration with path and dpi
    :return: Configured plot output formatter
    """
    if not isinstance(args, PlotBenchmarkOutputArgs):
        raise ValueError(f"Expected PlotBenchmarkOutputArgs, got {type(args)}")

    return cls(output_path=args.path, dpi=args.dpi)

GenerativeBenchmarkerSerialized

Bases: GenerativeBenchmarkerOutput

Serialized output handler for benchmark reports in JSON or YAML formats.

This output handler persists generative benchmark reports to the file system in either JSON or YAML format. It supports flexible path specification, allowing users to provide either a directory (where a default filename will be generated) or an explicit file path for the serialized report output.

Example: :: output = GenerativeBenchmarkerSerialized(output_path="/path/to/output.json") result_path = await output.finalize(report)

Source code in src/guidellm/benchmark/outputs/serialized.py
@GenerativeBenchmarkerOutput.register(["json", "yaml"])
class GenerativeBenchmarkerSerialized(GenerativeBenchmarkerOutput):
    """
    Serialized output handler for benchmark reports in JSON or YAML formats.

    This output handler persists generative benchmark reports to the file system in
    either JSON or YAML format. It supports flexible path specification, allowing
    users to provide either a directory (where a default filename will be generated)
    or an explicit file path for the serialized report output.

    Example:
    ::
        output = GenerativeBenchmarkerSerialized(output_path="/path/to/output.json")
        result_path = await output.finalize(report)
    """

    output_path: Path = Field(
        default_factory=Path.cwd,
        description="Directory or file path for saving the serialized report",
    )
    format_type: Literal["json", "yaml"] = Field(
        default="json",
        description="Serialization format, used to determine the default filename",
    )

    @classmethod
    def from_args(cls, args: BenchmarkOutputArgs) -> GenerativeBenchmarkerSerialized:
        """
        Create a serialized output formatter from output arguments.

        :param args: Output configuration with path and kind (json or yaml)
        :return: Configured serialized output formatter
        """
        if not isinstance(args, JSONBenchmarkOutputArgs | YAMLBenchmarkOutputArgs):
            raise ValueError(f"Invalid args type: {type(args)}.")

        return cls(output_path=args.path, format_type=args.kind)

    async def finalize(self, report: GenerativeBenchmarksReport) -> Path:
        """
        Serialize and save the benchmark report to the configured output path.

        :param report: The generative benchmarks report to serialize
        :return: Path to the saved report file
        """
        output_path = self.output_path
        if output_path.is_dir():
            output_path = output_path / f"benchmarks.{self.format_type}"
        return report.save_file(output_path)

finalize(report) async

Serialize and save the benchmark report to the configured output path.

Parameters:

Name Type Description Default
report GenerativeBenchmarksReport

The generative benchmarks report to serialize

required

Returns:

Type Description
Path

Path to the saved report file

Source code in src/guidellm/benchmark/outputs/serialized.py
async def finalize(self, report: GenerativeBenchmarksReport) -> Path:
    """
    Serialize and save the benchmark report to the configured output path.

    :param report: The generative benchmarks report to serialize
    :return: Path to the saved report file
    """
    output_path = self.output_path
    if output_path.is_dir():
        output_path = output_path / f"benchmarks.{self.format_type}"
    return report.save_file(output_path)

from_args(args) classmethod

Create a serialized output formatter from output arguments.

Parameters:

Name Type Description Default
args BenchmarkOutputArgs

Output configuration with path and kind (json or yaml)

required

Returns:

Type Description
GenerativeBenchmarkerSerialized

Configured serialized output formatter

Source code in src/guidellm/benchmark/outputs/serialized.py
@classmethod
def from_args(cls, args: BenchmarkOutputArgs) -> GenerativeBenchmarkerSerialized:
    """
    Create a serialized output formatter from output arguments.

    :param args: Output configuration with path and kind (json or yaml)
    :return: Configured serialized output formatter
    """
    if not isinstance(args, JSONBenchmarkOutputArgs | YAMLBenchmarkOutputArgs):
        raise ValueError(f"Invalid args type: {type(args)}.")

    return cls(output_path=args.path, format_type=args.kind)