Running Qwen3.8-27B locally on two RTX 3090s: vLLM, MTP and a private API
How I installed Qwen3.8-27B on a PC with two RTX 3090s: vLLM vs llama.cpp benchmarks, MTP, TP vs PP and an on-demand private API over Tailscale.
Development of the agent behind an inspection assistant for Red Eléctrica de España. Technicians had millions of photos of high-voltage towers and their defects, and had to review them tower by tower. With the agent they can query by area and by defect type, which makes prioritising and dispatching repair crews far more efficient.
Alvea hired me as a freelancer to build agent products with their partner company, Turing Dream. The main project was a multi-agent tutoring system for secondary school students, with one agent per subject and course material written and validated by a team of teachers. It was used with real students at a private school in Spain and at a secondary school in Colombia.
Development of a chatbot for HoReCa (hotels, restaurants and cafes), through which cooks and waiters can talk to the chatbot to make queries about recipes and menus.
Technical leadership on a generative AI initiative on Azure: architecture, technology selection and risk assessment. The programme was halted by the client.
Machine Learning Engineer, vision algorithms development for autonomous vehicle and leadership in RAG system to obtain documentation information
AI, HW and FW development.
HW and FW development.
Help project manager in project management. HW and FW development.
How I installed Qwen3.8-27B on a PC with two RTX 3090s: vLLM vs llama.cpp benchmarks, MTP, TP vs PP and an on-demand private API over Tailscale.
How an agent turns the name and arguments of a tool call into a local function, handles errors, and returns the result t...
How to accumulate arguments delivered as streaming deltas and why a tool must never run before the provider sends its te...
Let's talk.
maximofn@gmail.com
Machine Learning and AI specialist. I develop solutions with generative AI, intelligent agents and custom models.
AI agents, powered by LLMs, promise to transform applications. But are they simple executors today or future intelligent collaborators? To reach their...
Learn to create an IA system to execute efficiently on a device
Hugging Face spaces allow us to run models with very simple demos, but what if the demo breaks? Or if the user deletes it? That's why I've created docker containers with some interesting spaces, to be able to use them locally, whatever happens. In fact, if you click on any project view button, it may take you to a space that doesn't work.
Let's talk.
maximofn@gmail.com
Machine Learning and AI specialist. I develop solutions with generative AI, intelligent agents and custom models.
Dataset with jokes in English
Use: Fine-tuning text generation models for humor
Dataset with translations from English to Spanish
Use: Training English-Spanish translation models
Dataset with Netflix movies and series
Use: Netflix catalog analysis and recommendation systems