Describes a single language model available for inference, whether registered globally or scoped to a tenant. Instances are returned by LlmManager.getAvailableModels and LlmManager.getModelByName, and are supplied back to LlmManager.invokeInference (via InferenceParameters) to select which provider and model handle a request. The capabilities list records which features the model supports, using the CAP_ constants defined on this class.
Implements: Serializable
Properties
| Property | Returns | Description |
|---|---|---|
| capabilities | List<String> | The set of features this model supports, as a list of the CAP_ constant values defined on this class, for example CAP_VISION or CAP_STRUCTURED_OUTPUTS. Never null, but may be empty if no capabilities were declared. |
| maxContextTokens | int | The maximum combined size, in tokens, of the prompt and any conversation history the model can accept in a single inference request. |
| maxOutputTokens | int | The maximum number of tokens the model can generate in the response to a single inference request. |
| maxTokens | int | The maximum number of context tokens the model accepts, identical to getMaxContextTokens. |
| name | String | The identifying name of this model, as understood by its provider. Used to look up a model by name via LlmManager.getModelByName and to select which model an inference request runs against. |
| provider | String | The name of the LlmProvider that serves this model. LlmManager uses this to find the matching provider when invoking inference. |
Methods
getName()
Returns: String
The identifying name of this model, as understood by its provider. Used to look up a model by name via LlmManager.getModelByName and to select which model an inference request runs against.
getProvider()
Returns: String
The name of the LlmProvider that serves this model. LlmManager uses this to find the matching provider when invoking inference.
getMaxTokens()
Returns: int
The maximum number of context tokens the model accepts, identical to getMaxContextTokens.
getMaxContextTokens()
Returns: int
The maximum combined size, in tokens, of the prompt and any conversation history the model can accept in a single inference request.
getMaxOutputTokens()
Returns: int
The maximum number of tokens the model can generate in the response to a single inference request.
getCapabilities()
Returns: List<String>
The set of features this model supports, as a list of the CAP_ constant values defined on this class, for example CAP_VISION or CAP_STRUCTURED_OUTPUTS. Never null, but may be empty if no capabilities were declared.