9 min read
What is the Pyramid of Evidence Strength? (Part 2)
The Pyramid of Evidence Strength includes all types of scientific research, ranked by the strength of evidence of the results. Scientific research is divided into categories: expert opinions, case reports, in vitro, animal research, observational, experimental, systematic reviews and meta-analyses. These categories reflect the way the evidence came about, and the strength of the evidence determines the place in the pyramid.
This is a sequel to 'What is the Pyramid of Evidence Strength? (Part 1)'.

The Pyramid of Evidence Strength is not absolute in determining the strength of evidence of scientific research. The quality of a study is just as important. In part 2 of 'What is the Pyramid of Evidence Strength?' I go deeper into experimental research, (narrative versus systematic) reviews, meta-analyses and the quality of scientific research.
What is experimental research?
In experimental research, researchers control the environment and behaviour of the study participants as much as possible. They deliberately change one factor in the intervention group compared with the control group, while keeping other factors that may have an influence as much the same as possible. This makes it possible to look at whether the factor that was changed does or does not influence the outcome of the study.
Randomised controlled trials (RCTs) are the gold standard in science. RCTs are experimental in nature and have strict conditions that they must (and can) meet. This makes it possible to investigate a causal link between the intervention (for example healthy nutrition or individual nutrients) and the outcome (risk factors for diseases or the diseases themselves) of the study. The conditions for RCTs include several aspects that determine the final strength of the evidence of an RCT:
-
Randomisation: people are placed in the different groups within the study by chance. This gives everyone an equal chance of ending up in the control or intervention group. As a result, as many factors as possible (for example sex or lifestyle-related factors) are equally distributed over both groups and no bias arises.
-
Double blind: by blinding both the researcher and the participant, neither player in the study knows who is in which group. This means researchers cannot treat participants differently based on the group they are in, and participants do not start behaving differently because of knowing which group they are in. This is of course a very difficult condition in nutrition science: when people have to follow a certain eating pattern or eat a certain food, it is hard to blind the participants.
-
Cross-over: in a cross-over study, people move from the control group to the intervention group and vice versa. This way the effect of the intervention is not determined by a comparison between groups of different people, but the same people have been in both the control group and the intervention group. As a result, the effect of the intervention is based on the results within the same person. This makes it possible to rule out the influence of factors such as genetic predisposition.
-
Multicentre: this is when the study is carried out in several places (for example hospitals). As a result, people are recruited in different cities or countries and it can be checked whether that influences the effect of the intervention. For example, one hospital can be in a poorer part of a city/country compared with the other hospital, so the effect of the environment of the study itself is also taken into account.
-
Placebo: the control group receives the same as the intervention group, for example a pill. However, the pill of the control group contains no active substance. Placebo is an incredibly strong effect that can make people feel effects of a pill even though there is no active substance in it. Believing in it alone can already cause effects. There are even sham operations in which someone is not actually operated on, but still experiences positive effects from the sham operation. By giving the control group a placebo, the effect of the intervention cannot be attributed to placebo, when placebo does not show the same effect in the control group. This one is also difficult in nutrition science: there is always an active substance in the placebo.
Experimental research example
For example, researchers at Wageningen Universiteit set up the randomised, double-blind, multicentre, placebo-controlled study 'SOFA'. In this study the researchers randomly divided almost 600 people over two groups (randomised): the intervention group received 2 grams of fish oil daily (intervention) and the control group received a pill with 2 grams of sunflower oil (placebo, but sunflower oil also contains active substances such as fats, kcal and vitamins). The pills could not be told apart and the researchers also did not know who was in which group (double blind). The people were recruited from different cardiology clinics spread across Europe (multicentre). In total the people were followed for 12 months to see whether there was a difference in cardiovascular disease and deaths between the two groups.
There are many snags to RCTs focused on nutrition and health. The question is whether this gold standard is really that golden when it comes to nutrition. I will come back to this later in another part of this series.
What is a review?
A review is a literature study (secondary research) that gives an overview of the Body of Evidence on a specific topic. The literature is summarised and in some cases also assessed.
Roughly speaking, there are two different kinds of reviews: narrative and systematic. A narrative review gives an overview of the current literature on a certain topic. A narrative review usually contains a storyline in which several aspects of a topic are explained, such as the history, the mechanism and research with animals and humans. The authors of the review do not state how they found the literature they use or what the quality of the literature is. A systematic review does do this. In a systematic review, the authors are transparent about the method (such as keywords, inclusion and exclusion criteria and search strategy) that they used for searching and assessing the literature, after which the conclusion based on the literature is put in context. This makes the strength of evidence of a systematic review much stronger. This is mainly because you can follow how the authors arrived at their findings. With a narrative review you have no idea which studies were left out and why, so the risk of cherry picking is high.
Narrative review example
An example of a narrative review is a review by JJ, DiNicolantonio et al. that looks at the relationship between saturated fat, sugar and coronary heart disease. The authors use scientific research to tell about the background information, relationships between saturated fat and coronary heart disease, sugar and coronary heart disease, and the historical perspective, and they finish with a conclusion. However, they do not tell how they searched for literature, which criteria were used for including and excluding studies, and how they assessed the literature for quality. So how do you know whether they left out studies that go against the narrative? How do you know whether the studies that were included are of good quality? You have no idea -- and that is a problem.
Systematic review example
An example of a systematic review is a review by L, Hooper et al. that looks at the relationship between fats and cardiovascular disease. They describe exactly where and how they searched, how the studies were assessed, give a clear description of the studies, the limitations that come with the studies and which conclusion they draw from it.
The conclusions of the two reviews contradict each other. However, the review by Hooper is the only one where you can check how the conclusion came about. You could even repeat it yourself to see whether you get the same result. This makes the strength of evidence of Hooper stronger than that of DiNicolantonio. Precisely for this reason I prefer to place narrative reviews in the category 'Expert opinion', as I described in part 1 of the Pyramid of Evidence Strength.
What is a meta-analysis?
A meta-analysis combines the outcomes of primary research and carries out a new analysis, with a new overarching result as the outcome.
So Hooper and colleagues, on top of the systematic review, also carried out a meta-analysis. Meta-analyses go a step further than reviews and do not only summarise the literature but use the outcomes of the included studies to carry out a new analysis. In this way a new outcome is produced. With this, (smaller) studies can be combined to produce more power. In addition, it can be determined per study to what extent the result counts in the analysis, so studies of lower quality can count less strongly in the analysis. So Hooper combines all RCTs that looked at reducing saturated fat intake versus an eating pattern without change and the effect on mortality and cardiovascular disease. All outcomes of the studies are combined in the analysis and give an overarching result.
Quality and strength of evidence
The Pyramid of Evidence Strength was originally developed for clinical research focused on medication. As I said earlier, you cannot compare research on nutrition and medication 1-to-1. The effect of nutrition on our health is the result of years of exposure (chronic versus acute). In addition, other lifestyle-related factors also have an effect on the same outcomes. However, one of the biggest differences between nutrition and medication is that nutrition consists of tens of thousands of bioactive substances, which can all have an effect on our health and on each other (this is still a grey area). Plus, they are all also present in our body by default, because we eat every day. Medication is often about one or a few active substances that are often not present in our body, or only in small amounts. This makes nutrition science incredibly complex and the Pyramid of Evidence Strength not enough to assess studies: the quality of a study can be much more important than its place on the pyramid. A long-term prospective observational study can give a much clearer picture of the effect of nutrition on our health than a short-term RCT. That is why a systematic review assesses the studies it summarises for quality and design, which is crucial for being able to assess the strength of evidence.
Scientific research is assessed with methods that were developed for this. An important aspect of this is looking at the risk of bias. There are different methods for observational and experimental research, but they all look at the methodological design of the study and assess it for quality. The assessment of quality focuses on the influence of errors or flaws (which can be deliberate, accidental or unavoidable) in the design that can cause the results of the study to be influenced (bias). But important aspects a study can be assessed on also include whether a power analysis was done, how large the group of participants is and how representative the participants are of the larger population they were taken from. The authors of a study will always include these flaws in their discussion themselves to put the results in context, but authors will take care not to run down their own study. That is why, as a reader, you must always be aware of the possible flaws, and the writers of a systematic review assess every study they include for quality. Later in this series I will go deeper into the quality of studies.
Shit in Shit out
Even though reviews and meta-analyses are at the top of the pyramid, they are not always of good strength of evidence. The strength of evidence of a review and meta-analysis depends on the studies it contains. Shit in is shit out. A review and meta-analysis based on low-quality research has a low strength of evidence that does not suddenly go to the moon because the methods are at the top of the pyramid. Even though it is sometimes presented that way. That is why you must always look critically at the content, and preferably researchers do that themselves by writing a systematic review. This is just not always the case.
Determining the strength of evidence is therefore not as simple as the pyramid of evidence strength suggests. Several suggestions have also been made to adapt the pyramid. For example by making the layers of the pyramid more wavy. This makes it clear that, based on the quality of the study design, an observational study can have a higher strength of evidence than an experimental study. In addition, it has been proposed not to put reviews and meta-analyses at the top of the pyramid, but to use them as a magnifying glass. This makes it clear that reviews and the meta-analysis are instruments to summarise, assess and make applicable the results of other studies. This way no implicit quality is given to these instruments, but they are only as strong as the studies that are cited and assessed in them. Shit in shit out.
Hypothesis
The ultimate goal of that whole pyramid is to confirm a hypothesis. In other words, to find a significant result. More about hypothesis and significance in the next article 'What is a hypothesis?'.
What did you think? Let me know in the form of a comment or an email: [email protected]
Did you come across a claim on the internet or social media and are you curious about an assessment of its support, let me know and I will dive in!
Do you want to support me and my nutrition science adventure? Share my articles or the podcast on your socials!
Everything you read on this website is my opinion based on knowledge and experience. Chances are I sometimes overlook something, or that something could be better. I would love to hear it! We do science together.
Terms
Body of evidence: All results of all (kinds of) studies on a certain topic together. When we want to draw a conclusion about the effect of nutrition on our health, the evidence from all categories in the Pyramid of Evidence Strength must be combined into one story. In other words, the body of evidence.
Cherry picking: When studies or results of studies are chosen selectively and subjectively to support a claim, we speak of cherry picking.
