チャットの使用 / 埋め込みレスポンスの使用

概要

Spring AI は、使用状況インターフェースに getNativeUsage() メソッドを導入し、DefaultUsage 実装を提供することで、モデル使用状況の処理機能を強化しました。この変更により、フレームワーク全体の一貫性を維持しながら、さまざまな AI モデルが使用状況メトリクスを追跡および報告する方法が簡素化されます。

主な変更点

使用インターフェースの強化

Usage インターフェースに新しいメソッドが追加されました:

Object getNativeUsage();

このメソッドにより、モデル固有のネイティブ使用状況データにアクセスできるため、必要に応じてより詳細な使用状況追跡が可能になります。

ChatModel と併用

OpenAI の ChatModel を使用して使用状況を追跡する方法を示す完全な例を次に示します。

@SpringBootConfiguration
public class Configuration {

        @Bean
        public OpenAiChatModel openAiChatModel() {
            return OpenAiChatModel.builder()
                .options(OpenAiChatOptions.builder()
                    .apiKey(System.getenv("OPENAI_API_KEY"))
                    .build())
                .build();
        }

    }

@Service
public class ChatService {

    private final OpenAiChatModel chatModel;

    public ChatService(OpenAiChatModel chatModel) {
        this.chatModel = chatModel;
    }

    public void demonstrateUsage() {
        // Create a chat prompt
        Prompt prompt = new Prompt("What is the weather like today?");

        // Get the chat response
        ChatResponse response = this.chatModel.call(prompt);

        // Access the usage information
        Usage usage = response.getMetadata().getUsage();

        // Get standard usage metrics
        System.out.println("Prompt Tokens: " + usage.getPromptTokens());
        System.out.println("Completion Tokens: " + usage.getCompletionTokens());
        System.out.println("Total Tokens: " + usage.getTotalTokens());

        // Access native OpenAI usage data with detailed token information
        if (usage.getNativeUsage() instanceof com.openai.models.completions.CompletionUsage) {
            com.openai.models.completions.CompletionUsage nativeUsage =
                (com.openai.models.completions.CompletionUsage) usage.getNativeUsage();

            // Detailed prompt token information
            nativeUsage.promptTokensDetails().ifPresent(details -> {
                System.out.println("Prompt Tokens Details:");
                details.audioTokens().ifPresent(tokens -> System.out.println("- Audio Tokens: " + tokens));
                details.cachedTokens().ifPresent(tokens -> System.out.println("- Cached Tokens: " + tokens));
            });

            // Detailed completion token information
            nativeUsage.completionTokensDetails().ifPresent(details -> {
                System.out.println("Completion Tokens Details:");
                details.reasoningTokens().ifPresent(tokens -> System.out.println("- Reasoning Tokens: " + tokens));
                details.acceptedPredictionTokens().ifPresent(tokens -> System.out.println("- Accepted Prediction Tokens: " + tokens));
                details.audioTokens().ifPresent(tokens -> System.out.println("- Audio Tokens: " + tokens));
                details.rejectedPredictionTokens().ifPresent(tokens -> System.out.println("- Rejected Prediction Tokens: " + tokens));
            });
        }
    }
}

ChatClient と併用

ChatClient を使用している場合は、ChatResponse オブジェクトを使用して使用状況情報にアクセスできます。

// Create a chat prompt
Prompt prompt = new Prompt("What is the weather like today?");

// Create a chat client
ChatClient chatClient = ChatClient.create(chatModel);

// Get the chat response
ChatResponse response = chatClient.prompt(prompt)
        .call()
        .chatResponse();

// Access the usage information
Usage usage = response.getMetadata().getUsage();

プロンプトキャッシュ使用状況メトリクス

プロンプトキャッシュをサポートするプロバイダの場合、Usage インターフェースはプロバイダ固有の型変換を必要とせずにキャッシュメトリクスへの統一的なアクセスを提供します。

Usage usage = response.getMetadata().getUsage();

// Unified cache metrics — works across all providers
Long cacheReadTokens = usage.getCacheReadInputTokens();
Long cacheWriteTokens = usage.getCacheWriteInputTokens();

if (cacheReadTokens != null && cacheReadTokens > 0) {
    System.out.println("Cache hit: " + cacheReadTokens + " tokens read from cache");
}
if (cacheWriteTokens != null && cacheWriteTokens > 0) {
    System.out.println("Cache write: " + cacheWriteTokens + " tokens written to cache");
}

これらのメソッドは、プロンプトキャッシュをサポートしていないプロバイダに対して null を返します。

以下の表は、プロバイダ別のプロンプトキャッシュメトリクスの利用可能性を示しています。

プロバイダー キャッシュ読み取りトークン キャッシュ書き込みトークン

Anthropic

はい

はい (cacheCreationInputTokens)

AWS Bedrock

はい

はい

OpenAI

はい (cachedTokens)

いいえ

Google Gemini

はい (cachedContentTokenCount)

いいえ

DeepSeek

いいえ

いいえ

Mistral

いいえ

いいえ

Ollama

いいえ

いいえ

プロバイダー固有の詳細なキャッシュメトリクス(Gemini におけるモダリティごとのキャッシュ内訳など)については、getNativeUsage() を使用してプロバイダーのネイティブ使用状況オブジェクトにアクセスしてください。

複数ステップのフローにおける累積使用量

ツール呼び出しループなどの複数ステップのフローを通じてレスポンスが生成された場合、getUsage() は最後の呼び出しだけでなく、その交換におけるすべてのモデル呼び出しの累積トークン使用量を報告します。例: 1 つのツール呼び出しをトリガーする ChatClient 会話では、少なくとも 2 つのモデル呼び出しが実行され、返される getUsage() は両方の合計を反映します。

ChatResponse response = chatClient.prompt("What is the weather in Paris?")
        .tools(new WeatherTools())
        .call()
        .chatResponse();

// Cumulative across all model calls in the tool-calling loop
Usage usage = response.getMetadata().getUsage();
int totalTokens = usage.getTotalTokens();

累積合計は、標準トークン数と統合キャッシュメトリクス(getCacheReadInputTokens() / getCacheWriteInputTokens())を合計した org.springframework.ai.support.UsageCalculator によって計算されます。

プロバイダ固有のネイティブ使用オブジェクトは、複数のレスポンス間でマージすることはできません。そのため、getNativeUsage() は、複数のモデル呼び出しにわたって使用状況が集約された後 (たとえば、ツール呼び出しループの後) には null を返します。null は、単一呼び出しのレスポンスでのみ保持されます。プロバイダのネイティブ使用オブジェクトが必要な場合は、複数ステップの ChatClient 交換からではなく、個々の ChatModel 呼び出しから読み取ってください。

メリット

標準化 : さまざまな AI モデルの使用状況を一貫した方法で処理します。柔軟性 : ネイティブ使用状況機能簡素化を通じてモデル固有の使用状況データをサポートします: デフォルト実装拡張性で定型コードを削減: 互換性を維持しながら、特定のモデル要件に合わせて簡単に拡張できます

型安全性に関する考慮事項

ネイティブ使用状況データを扱うときは、型キャストを慎重に検討してください。

// Safe way to access native usage
if (usage.getNativeUsage() instanceof com.openai.models.completions.CompletionUsage) {
    com.openai.models.completions.CompletionUsage nativeUsage =
        (com.openai.models.completions.CompletionUsage) usage.getNativeUsage();
    // Work with native usage data
}