A Unifying Theory of the Evolution of Neurodiversity: Part 2
Reliable genomic instability as a clever trick to maintain diversity
I am applying for a postdoc. For prospective professors – this is one of the projects I could work on. A technical summary and possible research directions can be found at the end of this article.
Imagine that you’re a genome. And you want to be passed down from generation to generation. You desire many descendants that carry your genetic material.
Of course, as a genome, you don’t actually have will or desire. You accumulate random genetic variation, and the mutations that are consistently beneficial will undergo positive selection and eventually reach fixation at the population level. Mutations that are consistently detrimental will be eliminated from the population. Neutral mutations (that confer no benefit or detriment) will randomly fluctuate in frequency within the population and may reach fixation or be eliminated.
Not all mutations are random
But mutations in the genome aren’t completely random. There are some mutations that are frequently reintroduced into the human gene pool over and over and over again. For example, there is an unstable region on chromosome 16 that often experiences large duplication (and deletion) events. The resulting mutation is known as the 16p12.1 duplication (or deletion) and is associated with autism, intellectual disability, schizophrenia, and other developmental problems. I studied this mutation closely during my PhD work.
Large duplications and deletions in the genome are known as copy number variants, and copy number variants are strongly associated with neurodevelopmental disorders. Over the last three decades, studies have found an increased burden of de novo (as in ‘not inherited from the parents’ but ‘new’) copy number variants in individuals with autism, ADHD, schizophrenia and bipolar disorder. For each of these disorders, affected individuals have about a 5-fold increase in burden of de novo copy number variants than the general population. Some of the de novo copy number variants carried by these individuals are recurrent copy number variants and others are less common.
Now, I think intellectual disability is a confounder in autism studies and autism with and without intellectual disability should be studied separately. In this case, however, the association between de novo copy number variation and autism remains independent of intellectual disability.
We should expect genomic instabilities, if detrimental, to be eliminated from the genome. Why? Because having offspring is costly, and if a child is born with a genetic variant that confers reproductive disadvantages, it’s a huge spend of resources. Therefore, detrimental instability should be selected against just as much as the detrimental variants themselves.
However, I propose an explanation as to why instabilities may, in fact, be maintained in the genome. The hypothesis: Unstable regions in the genome are a “clever trick” to preserve diversity that is high cost to the individual but moderate to low benefit to the group.
This is a followup to my 2025 article on the same topic, which you can read here: https://aarysh.substack.com/p/the-ultimate-theory-on-the-evolution. In it, I propose a hypothesis that explains how common single nucleotide variants that confer risk towards neurodevelopmental disorders are maintained within the population. (As opposed to this article, in which I propose how recurrent copy number variants that confer risk towards neurodevelopmental disorders are maintained within the population.) Thanks a bunch to James Horton and the other commenters for their feedback.
(Reading Part 1 isn’t required to understand this article, but it does have a helpful explanation as to why neurodiversity phenotypes may be of benefit to the group, which I did not want to repeat here.)
Commenters swiftly pointed out that a related concept known as inclusive fitness was previously developed by evolutionary biologists to explain the evolution of traits that are costly to the individual but of benefit to the group. Inclusive fitness explains how a a single trait (like selfless altruism) is maintained within a population. On the other hand, I’m interested in explaining how a diversity of traits (like the diversity seen in the immune system or neuro traits) is maintained within a population. Nonetheless, the gene-centered view that inclusive fitness takes is interesting and useful.
One of the principles derived from work on inclusive fitness is Hamilton’s rule:
C < Br
For a trait that is costly to the individual to persist within a population, the reproductive cost (C) to the individual must be less than the reproductive benefit (B) to the individual’s relatives multiplied by the chance that the relative shares the same trait producing variant (r). So, for example, if I have a trait that makes it so I have 1 less child but I help my siblings (who share 50% of my DNA) have 2 extra children, then the inequality is satisfied and the trait may persist in the population.
My new proposal, unlike my 2025 proposal, takes a gene-centered view. Here, I propose that the genome may have evolved “reliable instability” to maintain traits that are high cost to the individual but mild to moderate benefit to relatives that share that instability. I think it’s possible to derive an inequality similar to Hamilton’s rule for this concept. My intuition tells me the new rule would look like this: Cd < Br. Where d is the chance of the recurrent variant arising de novo.
Colorblindness: a representative trait
Colorblindness is a relatively simple trait when it comes to its genetics. The genes that encode the red and green color receptors sit near each other on the X chromosome. If one of those color receptor genes breaks then that results in colorblindness.

In my previous article, I used colorblindness as a simple disease model to investigate the validity of the proposed hypothesis. In support of it, a study by Vertelli et. al. found evidence of balancing selection of common single nucleotide variants on the red color receptor gene in humans but not in chimpanzees.
Since writing that article, I’ve learned that colorblindness is sometimes reintroduced into the population through deletion and duplication events. So, colorblindness is a nice disease model for this article’s hypothesis, too!
In that same article, Vertelli et. al. searched for duplications and deletions of the color receptor genes within humans and within chimpanzees. In humans, they found that individuals carry between 1-5 copies of the green color receptor gene and about 5% of samples individuals carried a duplication of the red color receptor gene. In contrast, none of the 56 sequenced chimps had a duplicated red receptor gene. In a previous study they cite, none of 30 sequenced chimps carried a copy number variation in either red or green color receptor gene.
So, do humans and chimpanzees have the same rate of colorblindness introduced in the gene pool, and it’s just selected out in chimpanzees more quickly, or are humans more likely to experience duplication/deletion events in those genes? Vertelli et. al. attempt to answer that question in their article. My interpretation of their data (and their interpretation seems to agree, although I take a more direct approach) is that humans experience the duplication/deletion events more frequently than chimps do. Why? Because there were no chimps with duplications of either of these genes. Duplications are usually less harmful than deletions (with a deletion you lose a whole gene), and if duplications are less (or not at all) harmful, we would expect to see at least some of these chimps with duplications of either of these genes. We don’t, so I think the human mutation rate is elevated.
What does this mean: There may be something broken, or relaxed, within the human genome that makes it susceptible to duplication/deletion events—specifically at the red and green color receptor gene region but perhaps, too, in other parts of the genome.
What genes tend to be in regions of copy number instability?
This hypothesis begs the question: What genes are copy number unstable? We should expect that more critical genes, like housekeeping genes, less likely to be copy number unstable. A loss of a critical gene would be more damaging, so purifying selection should work more strongly to keep them safe. Do genes that confer benefit through diversity, like immunity genes, and, as I propose, neuro genes, more likely to be copy number unstable?
Well, there are several good studies that look at copy number variable regions. Regions may be copy number variable due to copy number instability or due to an ancient copy number variant that has been maintained in the population (for example, through balancing selection). Since the data is readily available, I’ll look at copy number variable regions.
In a 2006 overview, Redon et. al. identified regions in the human genome with copy number variability. They then asked the question: what biological functions are the genes in those regions involved in? They published the significant hits as seen in the table below. Those hits included: neurophysiological process, synaptogenesis, sensory perception, transmission of nerve impulse, and immune response.

Around the same time in 2006, Nguyen et. al. released a study that compared the copy number variable genes found within the human and mouse genomes. In the human genome, they, like Redon et. al., found an overabundance of genes involved in neurophysiological process, sensory perception, and immune response.
More strikingly, both Redon et. al. and Nguyen et. al. found an overabundance of olfactory receptor genes in human copy number variable regions. In mice, however, olfactory receptor genes were underrepresented within copy number variable regions. This is astonishing and supports the hypothesis put forth in this blog post. It suggests that humans benefit from diversity in olfactory receptor genes whereas mice are more intolerant to variation in olfactory receptor genes and have stricter purifying selection acting on those genes, and, possibly, on instabilities of those genes.


Furthermore, rhodopsin-like receptor activity, which is related to vision, was overrepresented in human copy number variant regions in Redon et. al.’s study. Whereas Ngyuen et. al. found an underrepresentation of rhodopsin-like receptor activity in copy number variable regions in mice. This is another striking difference between copy number variation in humans and mice.
Yet another 2006 study by Perry et. al. looked at the copy number variable regions in chimps. They found that, in chimps, copy number variable regions were enriched for immune response genes, but they did not find enrichment for any neuro related genes. They also performed an enrichment analysis of copy number variable genes shared between humans and chimps (copy number variable in both species) and again no neuro gene sets were enriched.


Other explanations for the existence of copy number unstable regions in the human genome
The literature frequently mentions the effect that copy number variation has on evolution. Greater copy number variation results in faster evolution—because greater variation provides more diversity evolution can act on. Review articles also consider the selection pressures on the copy number variants themselves, but not on the copy number instabilities that result in the copy number variants.
It’s well known that copy number variants are more likely to form in certain regions of the genome. For example, in repetitive regions of the genome, copy mistakes, like nonallelic homologous recombination, are more likely to occur that result in the introduction of de novo copy number variants. But what selection pressures do those regions experience? Perhaps such discussions do exist, but I just haven’t encountered them in my literature review or in my search.
Nevertheless, I think it’s still useful to review some of the previously discussed evolutionary mechanisms that copy number variants are involved in:
Neofunctionalization: When a gene is duplicated, the copy is free to independently accumulate new mutations and evolve novel functions. This is how red-green color vision is thought to have originally evolved in Old World primates. One of the color receptor genes (red or green) duplicated, and the duplicated copy evolved a different color sensitivity.
Gene dosage: Genes may adapt their copy number to different environments or for other reasons. For example, AMY is a gene that encodes a starch digestion enzyme and is present in greater copies in human populations that eat more starch. Adaptation to different environments is also one of the drivers of balancing selection.
Gene compensation: A duplicated gene may act as a protective buffer against deleterious mutations.
Permanent heterozygote: Some genes confer greater benefit when inherited in the heterozygote form. This is known as the heterozygote advantage. When a gene is duplicated on a single chromosome, it allows for a recombination event to occur such that a single chromosome carries both variants of that gene—thus creating a permanent heterozygote.
Multi-allelic diversification: In multi-allelic diversification, a duplicated gene becomes fixed in the population because it allows for greater diversity through the accumulation of new mutations or recombination events. This may be beneficial in places where balancing selection takes place, like the immune system genes.
Conclusions
Again, what I put forth is just a hypothesis. But I think a compelling one. Here, I proposed that instabilities in the genome could act as a mechanism to maintain diversity within a population. Furthermore, I proposed that neurodiversity phenotypes are maintained in the population, in part, due to these reliable instabilities.
The end.
Technical Summary: Instabilities in the genome exist such that recurrent copy number variants are frequently introduced into the population. Copy number variants, in general, and de novo copy number variants, in particular, are strongly associated with neurodevelopmental disorders. We should expect that instabilities that result in the recurrent introduction of variants that are high reproductive cost to the individual to be selected out of the population. However, I propose that reliable instabilities are a mechanism that allow variants that are high cost to the individual but mild to moderate benefit to the group to persist.
Possible research directions:
Computational modelling to see how such a system operates. Will diversity be maintained? What parameters and assumptions are required?
What is the relationship between reliable instability and balancing selection? Can I better elucidate it? Both mechanisms benefit greater diversity.
Are genes that are under greater selection pressures due to changing environments also more likely to be copy number unstable? More copy number variation means “faster evolution”. So, can frequent, changing evolutionary pressures result in greater copy number instability on the frequently pressured genes? For example, some domesticated species, like dogs, may face strong, changing selection pressures due to breeding. Do the selection pressures due to breeding create reliable instabilities that benefit “faster evolution” of certain genes. (This direction was suggested by my boyfriend, Calvin.)
I used copy number variable regions to investigate the hypothesis of this article. How well do copy number variable regions in the genome overlap with copy number unstable regions? Is there already a map of copy number unstable regions?
Are there any other traits (other than immunity and neuro) that benefit from diversity and these mechanisms?
Nonallelic homologous recombination due to presence of repetitive segments are a well known reason for copy number instability. What are other reasons? Why are the chimp color receptor genes copy number stable and not the human ones?
In eusocial insects like ants, behavioral differences between different workers seem to be driven by differences in gene expression. How do copy number variants affect gene expression? Do copy number variants generally affect only genes within the copy number variant region, or do they alter expression across the genome?
How does SNV diversity in the genome relate to CNV diversity? Do they overlap?

