IIIT-H Study Challenges ‘Bigger Is Better’ AI Theory for Brain Research

0
1

At the recently concluded International Conference on Machine Learning (ICML) in Seoul, IIIT-H researchers Prof. Bapi Raju and Vijay Rowtula presented their findings debunking the myth that larger language models understand the human brain better.

The ICML is globally renowned as one of the ‘Big Three’ machine learning conferences alongside NeurIPS (the annual conference on Neural Information Processing Systems) and ICLR (the international conference on learning representations). At the ICML, the world’s leading researchers converge to present and publish cutting-edge research on all aspects of machine learning used in closely related areas like artificial intelligence, statistics and data science, as well as important application areas such as machine vision, computational biology, speech recognition, robotics and neuroscience. Prof. Bapi Raju and PhD researcher Vijay Rowtula had a poster presentation for their work titled, “Linguistic Properties and Model Scale In Brain Encoding: From Small To Compressed Language Models”.
What They Did
“Since we cannot work directly on or probe brains, we do the next best thing: work on large language models that exhibit the same capabilities as that of the brain,” explains Vijay. Their research began on the premise first established by Richard Antonello’s Scaling Laws, published in the 2023 NeurIPS paper, that LLMs tend to represent language in ways that more closely predict human brain activity and that moving from small to large models improved prediction accuracy by 15%. The IIIT-H researchers however tried to disprove Antonello’s findings and instead worked on small language models to demonstrate that they perform as comparably as LLMs. They in fact went a step further and tried to make SLMs smaller by introducing compression techniques such as quantization and pruning.
Quantization reduces the numerical precision used to represent information inside a model. Pruning, meanwhile, removes parts of the model considered less important. Both approaches can reduce the memory and computational resources required to run an AI system. “It’s like severing neurons and making them more lightweight. The novelty in our research was that we used multiple quantization techniques and then used frameworks to test the behaviours of these models on language,” elaborates Vijay.
What They Found
Most of the compression techniques tested including quantization and moderate pruning reduced model size without significantly harming the model’s ability to predict brain activity. The researchers found that models with about 3 billion parameters performed almost as well as much larger models with up to 14 billion parameters. The study also revealed that compression could weaken a model’s abilities on language tasks involving grammar, discourse and morphology, while its alignment with brain activity remained largely unchanged. “The novel finding is the dissociation observed between brain alignment and linguistic competence of language models – small as well as the compressed versions. It shows that it is possible that the capabilities needed to perform well on conventional language benchmarks may not exactly be the same as the representations that matter when it comes to modelling how the human brain processes language,” states Prof. Raju.
The broader implication is that compact AI models may be sufficient for studying language processing in the brain. Such models could make computational neuroscience research cheaper, faster and easier to conduct, without requiring researchers to work with the largest models available. According to the professor, “This might be a game changer for brain decoding workflows important for designing brain computer interfaces.”
From Computer Vision To Computational Neuroscience
For Vijay, the work also represents a new direction in an academic journey that has taken an unconventional route. He completed his Master’s at IIIT-H in 2019 under Prof. C.V. Jawahar, working in computer vision. He then spent several years in industry, including as a principal researcher, before returning to academia to pursue a PhD. in computational neuroscience.
The move from computer vision to studying language and the brain may appear like a significant shift. But Vijay sees the broader connection in the underlying computational models. “What I’m doing is computational neuroscience,” he says. “It requires imitating the behaviour of the brain using computational models.” His work is also connected to a larger question which is at the heart of the current AI race: how close can artificial systems come to human intelligence?
“Can a computational model imitate a human brain?” he asks. “The whole idea of my PhD is to contribute to the current AI race where we try to improve AI to the human level.” As AI systems increasingly move across domains, the distinctions between language, vision and other forms of intelligence are also becoming less clear. “There’s a fine line between a language model and a vision model. Under the hood, it is always the same transformer,” Vijay says. His return to academia has allowed him to explore these questions from a different vantage point, bringing with him the experience of working in industry coupled with a prior background in computer vision.

Disclaimer : This story is auto aggregated by a computer programme and has not been created or edited by DOWNTHENEWS. Publisher: deccanchronicle.com