Can minimal individual neural controllers, evolved under selection for collective performance, produce emergent division of labor in an artificial ant colony? We investigate this question through Evo-ACO, a neuroevolutionary extension of the Max–Min Ant System (MMAS) for the Traveling Salesman Problem. Every ant carries a lightweight micro-neural network (10 trainable weights, no hidden layer) that modulates its pheromone sensitivity and exploration probability at each construction step based on four real-time context signals. Weights are optimized online through steady-state neuroevolution with tournament selection, Gaussian mutation with adaptive sigma decay, and elitism. Analysis of evolved weights reveals emergent behavioral specialization: without any explicit diversity mechanism, ants self-organize into exploiter, explorer, and balanced roles, paralleling the response threshold model of division of labor in social insects. This emergent heterogeneity also yields strong optimization performance: across 52 TSPLIB instances (from 14 to 1084 vertices), Evo-ACO achieves (i) an average optimality gap of 0.32%; (ii) a 48% reduction over MMAS (0.62%), and (iii) compares favorably with reported results from AACO-LST, LBSA-CO, QLGLS, POMO and LR+P. Further, on 48 additional larger instances (with up to 18 512 vertices), it outperforms MMAS on 44/48 and compares favorably to QLGLS results on all 26 common instances. Issue Section:

Evolved Micro-Neural Networks for Per-Ant Behavioral Modulation in Ant Colony Optimization

Alessio Mezzina
Primo
;
Mario Pavone
2026-01-01

Abstract

Can minimal individual neural controllers, evolved under selection for collective performance, produce emergent division of labor in an artificial ant colony? We investigate this question through Evo-ACO, a neuroevolutionary extension of the Max–Min Ant System (MMAS) for the Traveling Salesman Problem. Every ant carries a lightweight micro-neural network (10 trainable weights, no hidden layer) that modulates its pheromone sensitivity and exploration probability at each construction step based on four real-time context signals. Weights are optimized online through steady-state neuroevolution with tournament selection, Gaussian mutation with adaptive sigma decay, and elitism. Analysis of evolved weights reveals emergent behavioral specialization: without any explicit diversity mechanism, ants self-organize into exploiter, explorer, and balanced roles, paralleling the response threshold model of division of labor in social insects. This emergent heterogeneity also yields strong optimization performance: across 52 TSPLIB instances (from 14 to 1084 vertices), Evo-ACO achieves (i) an average optimality gap of 0.32%; (ii) a 48% reduction over MMAS (0.62%), and (iii) compares favorably with reported results from AACO-LST, LBSA-CO, QLGLS, POMO and LR+P. Further, on 48 additional larger instances (with up to 18 512 vertices), it outperforms MMAS on 44/48 and compares favorably to QLGLS results on all 26 common instances. Issue Section:
File in questo prodotto:
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.11769/733610
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
social impact