NoticiasNews

Alibaba lanza Qwen-UI-Agent y apunta al software sin API: los sistemas viejos también se pueden automatizarAlibaba Launches Qwen-UI-Agent and Targets Software Without APIs: Legacy Systems Can Be Automated Too

2026-08-25

El 22 de agosto de 2026 Alibaba presentó Qwen-UI-Agent, un modelo base orientado a interfaces gráficas capaz de operar teléfonos, computadoras, aplicaciones web y entornos de búsqueda profunda entendiendo directamente los elementos en pantalla y ejecutando clics, acciones y tareas de varios pasos. Según la compañía, en múltiples benchmarks reconocidos de GUI supera a modelos de primera línea como GPT-5.6 y Claude Opus 4.8, lo que sugiere mayor fiabilidad en navegación de interfaces y finalización de tareas.

Por qué esto importa más en LATAM que en Silicon Valley

Buena parte de los casos de uso de agentes se rompen cuando el modelo tiene que operar software real de escritorio o móvil en vez de llamar a una API limpia. Y ese es exactamente el escenario dominante en la región. Un estudio de Seren junto con la firma de análisis PCMI, citado durante el Banking Tech Summit celebrado en República Dominicana en 2026, estima que alrededor del 60% de las entidades bancarias de la región sigue operando con núcleos heredados de más de 30 años. A eso se suman los sistemas administrativos, contables y operativos de miles de empresas medianas que nunca expusieron un endpoint porque nunca hizo falta.

Hasta ahora, automatizar sobre ese tipo de software implicaba una de tres cosas: reemplazarlo (caro y lento), envolverlo en una capa de integración a medida (viable pero con esfuerzo), o resignarse a que una persona siga haciendo clics. Un agente base sólido en GUI convierte las tareas basadas en pantalla —flujos tipo RPA, navegación en aplicaciones empresariales, herramientas legadas sin API— en objetivos de automatización de primera clase en lugar de casos límite.

La advertencia que viene con la promesa

Conviene leer el anuncio con la cabeza fría. Los agentes que operan interfaces fallan de una manera particularmente incómoda: fallan en silencio. Un clic en el botón equivocado no lanza una excepción ni deja un error en el log; simplemente produce un resultado incorrecto que nadie nota hasta que se acumula. Esto es distinto a una integración por API, donde un contrato roto se manifiesta de inmediato.

Por eso, quien esté evaluando este tipo de automatización debería:

  • Empezar por dos o tres procesos repetitivos que hoy dependen de que una persona navegue una interfaz compleja, y medir tasa de éxito y perfil de errores antes de ampliar.
  • Diseñar rutas explícitas de escalamiento en vez de asumir autonomía perfecta: qué hace el agente cuando no está seguro.
  • Instrumentar verificación posterior, es decir, validar el resultado del proceso —no solo que el agente terminó—.
  • Definir qué pantallas quedan prohibidas, especialmente aquellas donde una acción es irreversible o mueve dinero.

Automatizar la pantalla no elimina la deuda técnica

Hay una tentación evidente: si un agente puede operar el sistema viejo, ya no hace falta modernizarlo. Es una conclusión peligrosa. Un agente sobre una interfaz legada es una solución de puente, no de arquitectura. Alivia el trabajo manual, pero no resuelve la fragilidad, ni la falta de datos estructurados, ni la imposibilidad de auditar lo que ocurrió por dentro. La estrategia sensata combina las dos cosas: automatizar la pantalla para ganar tiempo hoy, y usar ese tiempo para construir las integraciones que el negocio va a necesitar mañana.

Dónde entra TEKFENIX

En TEKFENIX trabajamos precisamente en esa costura entre lo que ya existe y lo que hace falta. Desarrollamos software empresarial a medida e integraciones que convierten sistemas cerrados en procesos consultables y auditables. Servigo365 centraliza la atención multicanal con IA sin obligarle a reemplazar sus herramientas actuales; Nexturno ordena los turnos y colas en sucursal conectándose a la operación real; y CumplimientoControl asegura que la trazabilidad regulatoria no se pierda cuando parte del proceso pasa a ejecutarse de forma automática. Si tiene un sistema que nadie quiere tocar pero que todos usan a diario, ese es exactamente el problema que sabemos resolver.

On August 22, 2026, Alibaba introduced Qwen-UI-Agent, a GUI-focused base agent able to operate phones, PCs, web apps and deep search environments by directly understanding on-screen elements and executing clicks, actions and multi-step tasks. According to the company, it outperforms flagship models such as GPT-5.6 and Claude Opus 4.8 on multiple recognized GUI benchmarks, suggesting stronger reliability in interface navigation and task completion.

Why this matters more in LATAM than in Silicon Valley

Many agent use cases break down when the model must operate real desktop or mobile software instead of calling a clean API. That is precisely the dominant scenario in the region. A study by Seren with analysis firm PCMI, cited during the Banking Tech Summit held in the Dominican Republic in 2026, estimates that roughly 60% of banking entities in the region still run legacy cores over 30 years old. Add to that the administrative, accounting and operational systems of thousands of mid-sized companies that never exposed an endpoint because they never needed to.

Until now, automating on top of that kind of software meant one of three things: replacing it (expensive and slow), wrapping it in a custom integration layer (viable but effortful), or accepting that a person keeps clicking. A solid GUI-capable base agent turns screen-based tasks — RPA-style flows, enterprise app navigation, legacy tools with no API — into first-class automation targets instead of edge cases.

The warning that comes with the promise

It is worth reading the announcement with a cool head. Agents that operate interfaces fail in a particularly uncomfortable way: they fail silently. A click on the wrong button throws no exception and leaves no error in the log; it simply produces an incorrect result nobody notices until it accumulates. This differs from an API integration, where a broken contract shows up immediately.

Anyone evaluating this kind of automation should therefore:

  • Start with two or three repetitive processes that today depend on a person navigating a complex interface, and measure success rate and error profile before expanding.
  • Design explicit escalation paths instead of assuming perfect autonomy: what the agent does when it is not sure.
  • Instrument downstream verification — validate the process outcome, not just that the agent finished.
  • Define which screens are off limits, especially where an action is irreversible or moves money.

Automating the screen does not erase technical debt

There is an obvious temptation: if an agent can operate the old system, there is no need to modernize it. That is a dangerous conclusion. An agent on a legacy interface is a bridge, not an architecture. It relieves manual work but does not fix fragility, the lack of structured data, or the inability to audit what happened underneath. The sensible strategy combines both: automate the screen to buy time today, and use that time to build the integrations the business will need tomorrow.

Where TEKFENIX fits

At TEKFENIX we work exactly on that seam between what already exists and what is missing. We build custom enterprise software and integrations that turn closed systems into queryable, auditable processes. Servigo365 centralizes AI-assisted multichannel support without forcing you to replace your current tools; Nexturno organizes branch queues and turns while connecting to real operations; and CumplimientoControl ensures regulatory traceability is not lost when part of the process starts running automatically. If you have a system nobody wants to touch but everybody uses daily, that is exactly the problem we know how to solve.

← Volver al blog← Back to blog