Bases: DatasetPreprocessor
Extract user prompts, system prompts, and tool responses from messages.
Many tool calling datasets (e.g. madroid/glaive-function-calling-openai) store conversations as a messages column containing an array of message dicts. This preprocessor supports both OpenAI format (role/content) and ShareGPT format (from/value).
It replaces the text_column value with the extracted user content, populates prefix_column with the system prompt when present, and populates tool_response_column with tool response content.
Usage::
guidellm benchmark run \
--data '{"kind": "hf", "source": "..."}' \
--data-column-mapper '{"kind": "generative_column_mapper",
"column_mappings": {"text_column": "messages"}}' \
--data-preprocessor kind=tool_calling_message_extractor
Source code in src/guidellm/data/preprocessors/tool_calling.py
| @PreprocessorRegistry.register("tool_calling_message_extractor")
class ToolCallingMessageExtractor(DatasetPreprocessor):
"""Extract user prompts, system prompts, and tool responses from messages.
Many tool calling datasets (e.g. ``madroid/glaive-function-calling-openai``)
store conversations as a ``messages`` column containing an array of
message dicts. This preprocessor supports both OpenAI format
(``role``/``content``) and ShareGPT format (``from``/``value``).
It replaces the ``text_column`` value with the extracted user content,
populates ``prefix_column`` with the system prompt when present, and
populates ``tool_response_column`` with tool response content.
Usage::
guidellm benchmark run \\
--data '{"kind": "hf", "source": "..."}' \\
--data-column-mapper '{"kind": "generative_column_mapper",
"column_mappings": {"text_column": "messages"}}' \\
--data-preprocessor kind=tool_calling_message_extractor
"""
def __init__(self, config: ToolCallingMessageExtractorArgs, **_: Any) -> None:
pass
def __call__( # noqa: C901
self, items: list[dict[str, Any]]
) -> list[dict[str, Any]]:
for item in items:
text_values = item.get("text_column")
if not text_values or not isinstance(text_values, list):
continue
new_texts: list[str] = []
prefixes: list[str] = []
tool_responses: list[str] = []
for value in text_values:
if isinstance(value, list):
user_parts, system_parts, tool_parts = _extract_from_messages(value)
if user_parts:
new_texts.append(" ".join(user_parts))
if system_parts:
prefixes.append(" ".join(system_parts))
tool_responses.extend(tool_parts)
elif isinstance(value, str):
new_texts.append(value)
if new_texts:
item["text_column"] = new_texts
if prefixes:
item.setdefault("prefix_column", []).extend(prefixes)
if tool_responses:
item.setdefault("tool_response_column", []).extend(tool_responses)
return items
|