What happened
OpenAIEmbeddingModel keys its embedding cache only by the API request parameters. The endpoint is missing from that identifier. Two OpenAI-compatible endpoints using the same model name, dimensions and input therefore share a cache entry, even when they serve different embedding models.
With a shared FileEmbeddingCache, changing the endpoint returns the first endpoint's vectors without making a request to the second endpoint. This can silently mix embedding spaces when switching a RAG backend.
Reproduction
Tested on main at 72f3f6fa0b2fc38b8517f408ab616f0f2bd229e6, AgentScope 2.0.10.dev0, Python 3.14.3 and OpenAI 3.24.0. The check uses the real OpenAI SDK and two local HTTP servers; no provider credentials or external API calls are needed.
The two servers both expose /v1/embeddings and accept model="shared-name", dimensions=2, input=["hello"]. Server A returns [[1.0, 0.0]]; server B returns [[0.0, 1.0]].
cache = FileEmbeddingCache(cache_dir)
model_a = OpenAIEmbeddingModel(
credential=OpenAICredential(api_key="local-test-key", base_url=url_a),
model="shared-name", dimensions=2, embedding_cache=cache,
)
model_b = OpenAIEmbeddingModel(
credential=OpenAICredential(api_key="local-test-key", base_url=url_b),
model="shared-name", dimensions=2, embedding_cache=cache,
)
first = await model_a(["hello"])
second = await model_b(["hello"])
Actual result:
first: embeddings=[[1.0, 0.0]], source='api'
second: embeddings=[[1.0, 0.0]], source='cache'
server A requests: 1
server B requests: 0
Expected: the second call reaches server B and returns [[0.0, 1.0]]. A subsequent call to the same endpoint should still hit its cache.
Proposed fix
Include the SDK's resolved client.base_url in the cache identifier, using a separate identifier from the API kwargs. This also accounts for the SDK's default/environment-provided URL and trailing-slash normalization. API request parameters should stay unchanged, and API keys should not be part of the identifier.
Existing endpoint-agnostic entries cannot safely be attributed to a provider, so treating them as cache misses is preferable to reusing potentially wrong vectors. I can submit a focused fix with endpoint isolation and same-endpoint reuse tests.
What happened
OpenAIEmbeddingModelkeys its embedding cache only by the API request parameters. The endpoint is missing from that identifier. Two OpenAI-compatible endpoints using the same model name, dimensions and input therefore share a cache entry, even when they serve different embedding models.With a shared
FileEmbeddingCache, changing the endpoint returns the first endpoint's vectors without making a request to the second endpoint. This can silently mix embedding spaces when switching a RAG backend.Reproduction
Tested on
mainat72f3f6fa0b2fc38b8517f408ab616f0f2bd229e6, AgentScope2.0.10.dev0, Python 3.14.3 and OpenAI 3.24.0. The check uses the real OpenAI SDK and two local HTTP servers; no provider credentials or external API calls are needed.The two servers both expose
/v1/embeddingsand acceptmodel="shared-name",dimensions=2,input=["hello"]. Server A returns[[1.0, 0.0]]; server B returns[[0.0, 1.0]].Actual result:
Expected: the second call reaches server B and returns
[[0.0, 1.0]]. A subsequent call to the same endpoint should still hit its cache.Proposed fix
Include the SDK's resolved
client.base_urlin the cache identifier, using a separate identifier from the API kwargs. This also accounts for the SDK's default/environment-provided URL and trailing-slash normalization. API request parameters should stay unchanged, and API keys should not be part of the identifier.Existing endpoint-agnostic entries cannot safely be attributed to a provider, so treating them as cache misses is preferable to reusing potentially wrong vectors. I can submit a focused fix with endpoint isolation and same-endpoint reuse tests.