> For the complete documentation index, see [llms.txt](https://langtech-bsc.gitbook.io/alia-kit/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://langtech-bsc.gitbook.io/alia-kit/modelos/modelos-multimodales.md).

# Modelos multimodales

<table data-view="cards"><thead><tr><th>Descripción / Función</th><th>Nombre modelo</th><th data-type="content-ref">Model card</th><th>Comentario</th></tr></thead><tbody><tr><td>LLM especializado en imágenes y videos</td><td>Salamandra-VL-7B-2512</td><td><a href="https://huggingface.co/BSC-LT/Salamandra-VL-7B-2512">https://huggingface.co/BSC-LT/Salamandra-VL-7B-2512</a></td><td>Versión más reciente de la familia de modelos multimodales Salamandra. Combina el codificador visual <a href="https://huggingface.co/google/siglip2-giant-opt-patch16-384">SigLIP 2 Giant</a> con <a href="https://huggingface.co/BSC-LT/salamandra-7b">Salamandra 7B</a>  ajustado para seguir instrucciones, con especial atención a las lenguas europeas, y mejora la comprensión visual detallada y el conteo mediante datos <a href="https://huggingface.co/collections/allenai/pixmo">PixMo</a>.</td></tr><tr><td>LLM especializado en imágenes y videos</td><td>salamandra-7b-vision</td><td><a href="https://huggingface.co/BSC-LT/salamandra-7b-vision">https://huggingface.co/BSC-LT/salamandra-7b-vision</a></td><td>Modelo salamandra-7b adaptado para el procesamiento de imágenes y videos.</td></tr><tr><td>Traducción de voz a texto</td><td>SalamandraTAV-7b</td><td><a href="https://huggingface.co/BSC-LT/salamandra-TAV-7b">https://huggingface.co/BSC-LT/salamandra-TAV-7b</a></td><td>Modelo de lenguaje multimodal especializado en voz y traducción, afinado a partir de <a href="https://huggingface.co/BSC-LT/salamandraTA-7b-instruct">salamandraTA-7b-instruct</a>, Admite seis lenguas ibéricas además del inglés y puede realizar reconocimiento automático del habla, traducción de texto, traducción de voz a texto e identificación de la lengua hablada.</td></tr><tr><td>Modelo multimodal (voz y texto)</td><td>speech-salamandra-es-en</td><td><a href="https://huggingface.co/BSC-LT/speech-salamandra-es-en">https://huggingface.co/BSC-LT/speech-salamandra-es-en</a></td><td>Modelo multimodal de voz y texto en español e inglés, basado en Salamandra-7B-Instruct e integrado con el codificador de voz de SeamlessM4T v2. Puede recibir instrucciones de texto y audio para realizar reconocimiento automático del habla y responder preguntas sobre contenidos hablados.</td></tr><tr><td>Modelo multimodal y muiltilingüe instruido</td><td>Latxa-Qwen3.5-2B</td><td><a href="https://huggingface.co/HiTZ/Latxa-Qwen3.5-2B">https://huggingface.co/HiTZ/Latxa-Qwen3.5-2B</a></td><td>Modelo multimodal y multilingüe instruido, basado en Qwen3-VL-2B-Instruct, capaz de comprender y generar texto, procesar imágenes y seguir instrucciones. Ha sido adaptado para mejorar su rendimiento en euskera y, en su variante multilingüe, también en gallego y catalán.</td></tr><tr><td>Modelo multimodal y muiltilingüe instruido</td><td>Latxa-Qwen3.5-4B</td><td><a href="https://huggingface.co/HiTZ/Latxa-Qwen3.5-4B">https://huggingface.co/HiTZ/Latxa-Qwen3.5-4B</a></td><td>Modelo multimodal y multilingüe instruido, basado en Qwen3.5-4B, capaz de comprender y generar texto, procesar imágenes y seguir instrucciones. Ha sido adaptado para mejorar su rendimiento en euskera, gallego y catalán.</td></tr><tr><td>Modelo multimodal y muiltilingüe instruido</td><td>Latxa Qwen-3 VL 2B</td><td><a href="https://huggingface.co/HiTZ/Latxa-Qwen3-VL-2B-Instruct">https://huggingface.co/HiTZ/Latxa-Qwen3-VL-2B-Instruct</a></td><td>Modelo multimodal y multilingüe instruido, basado en <a href="https://huggingface.co/Qwen/Qwen3-VL-2B-Instruct">Qwen3-VL-2B-Instruct</a>, capaz de comprender y generar texto, procesar imágenes y seguir instrucciones. Ha sido adaptado para mejorar su rendimiento en euskera y, en su variante multilingüe, también en gallego y catalán.   </td></tr><tr><td>Modelo multimodal y muiltilingüe instruido</td><td>Latxa Qwen-3 VL 4B</td><td><a href="https://huggingface.co/HiTZ/Latxa-Qwen3-VL-4B-Instruct">https://huggingface.co/HiTZ/Latxa-Qwen3-VL-4B-Instruct</a></td><td>Modelo multimodal y multilingüe instruido, basado en <a href="https://huggingface.co/Qwen/Qwen3-VL-4B-Instruct">Qwen3-VL-4B-Instruct</a>, capaz de comprender y generar texto, procesar imágenes y seguir instrucciones. Ha sido adaptado para mejorar su rendimiento en euskera y, en su variante multilingüe, también en gallego y catalán.</td></tr><tr><td>Modelo multimodal y muiltilingüe instruido</td><td>Latxa-Qwen3-VL-8B-Instruct</td><td><a href="https://huggingface.co/HiTZ/Latxa-Qwen3-VL-8B-Instruct">https://huggingface.co/HiTZ/Latxa-Qwen3-VL-8B-Instruct</a></td><td>Modelo multimodal y multilingüe instruido, basado en <a href="https://huggingface.co/Qwen/Qwen3-VL-8B-Instruct">Qwen3-VL-8B-Instruct</a>, capaz de comprender y generar texto, procesar imágenes y seguir instrucciones. Ha sido adaptado para mejorar su rendimiento en euskera y, en su variante multilingüe, también en gallego y catalán.</td></tr><tr><td>Modelo multimodal y muiltilingüe instruido</td><td>Latxa-Qwen3-VL-32B-Instruct</td><td><a href="https://huggingface.co/HiTZ/Latxa-Qwen3-VL-32B-Instruct">https://huggingface.co/HiTZ/Latxa-Qwen3-VL-32B-Instruct</a></td><td>Modelo multimodal y multilingüe instruido, basado en <a href="https://huggingface.co/Qwen/Qwen3-VL-32B-Instruct">Qwen3-VL-32B-Instruct</a>, capaz de comprender y generar texto, procesar imágenes y seguir instrucciones. Ha sido adaptado para mejorar su rendimiento en euskera y, en su variante multilingüe, también en gallego y catalán.</td></tr></tbody></table>


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://langtech-bsc.gitbook.io/alia-kit/modelos/modelos-multimodales.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
