Overview
PeptideInsight uses a structured evidence grading system to help readers quickly assess the quality and quantity of scientific research behind each peptide and its claimed effects. This system is grounded in the principles of evidence-based medicine (EBM) and adapted specifically for the peptide research landscape, where much of the available data is preclinical.
Evidence grades are assigned per peptide and, where applicable, per specific claimed effect. A single peptide may have strong evidence for one application, moderate evidence for another, and only preclinical data for a third. This granularity is important because marketing materials and online discussions frequently generalize from a peptide's strongest-supported use to all of its claimed benefits.
The Evidence Hierarchy
Before explaining our specific grading system, it is important to understand the broader hierarchy of scientific evidence that underpins it. Listed from strongest to weakest:
Level 1: Systematic Reviews and Meta-Analyses
A systematic review uses a predefined, reproducible search strategy to identify all relevant studies on a question, then critically appraises and synthesizes their findings. A meta-analysis goes further by statistically combining the quantitative results from multiple studies to produce a pooled effect estimate with greater precision than any individual study.
Why this ranks highest: By aggregating data across multiple trials, meta-analyses increase statistical power, can detect small but real effects, and help identify consistency (or inconsistency) of results across different populations and settings.
Caveats: A meta-analysis is only as good as the studies it includes. Pooling flawed studies produces a flawed synthesis — "garbage in, garbage out." Heterogeneity between studies (different populations, dosing, endpoints) can make pooling inappropriate.
Level 2: Randomized Controlled Trials (RCTs)
Participants are randomly assigned to receive either the treatment or a control (placebo or active comparator). Randomization minimizes selection bias and balances known and unknown confounders between groups. Double-blinding (neither participant nor investigator knows the assignment) further reduces measurement and reporting bias.
Why this ranks highly: Randomization is the most reliable method for establishing causation. If the only systematic difference between groups is the treatment, then differences in outcomes can be attributed to the treatment.
Key quality factors:
- Sample size and statistical power
- Adequacy of randomization and allocation concealment
- Blinding (single, double, or open-label)
- Intention-to-treat analysis
- Predefined primary endpoints
- Dropout rate and handling of missing data
- Follow-up duration
Level 3: Cohort Studies
Observational studies that follow groups of people over time, comparing those exposed to a factor (or treatment) with those not exposed. Prospective cohort studies follow participants forward in time; retrospective cohorts look back at historical data.
Strengths: Can study long-term outcomes, rare exposures, and multiple outcomes simultaneously. Useful when RCTs would be unethical or impractical.
Weaknesses: Cannot establish causation due to confounding variables. People who choose to take a peptide may differ systematically from those who do not.
Level 4: Case-Control Studies
Researchers identify people with an outcome (cases) and without (controls), then look backward to compare exposures. Useful for studying rare diseases.
Weaknesses: Highly susceptible to recall bias and selection bias. Cannot establish causation.
Level 5: Case Series and Case Reports
Descriptions of individual patients or small groups of patients. They can generate hypotheses and identify rare adverse events but cannot establish efficacy or causation.
Level 6: Animal Studies (In Vivo)
Studies conducted in living animals, most commonly mice and rats. These provide essential information about mechanisms, toxicity, and pharmacokinetics in a complete biological system, but results frequently fail to translate to humans. Approximately 90% of drugs that succeed in animal studies fail in human clinical trials.
Level 7: In Vitro Studies
Experiments conducted in cell cultures, tissue preparations, or isolated biochemical systems. Valuable for understanding molecular mechanisms but the most removed from clinical relevance.
Level 8: Expert Opinion and Mechanistic Reasoning
Theoretical arguments based on known biology or expert consensus without direct empirical testing. The weakest form of evidence, though sometimes the only form available for novel compounds.
Our Four Evidence Levels
Based on the hierarchy above, PeptideInsight assigns one of four evidence grades (plus an "Insufficient Data" category) to each peptide-indication pair.
Strong Evidence
Criteria — all of the following must be met:
- At least two well-designed RCTs in humans with consistent positive results
- Preferably supported by at least one systematic review or meta-analysis
- Results replicated by independent research groups (not all from one lab)
- Published in peer-reviewed journals with established impact factors
- Clinically meaningful effect sizes (not just statistically significant)
- Typically has at least one FDA, EMA, or other major regulatory agency approval
What this means for readers: The evidence supports that this peptide works for this specific indication in humans, with known safety parameters. This does not mean the peptide is risk-free or appropriate for all patients.
Examples:
- Semaglutide for type 2 diabetes and weight management (dozens of Phase III RCTs, multiple meta-analyses)
- Tirzepatide for type 2 diabetes and obesity (SURPASS and SURMOUNT trial programs)
- Bremelanotide for hypoactive sexual desire disorder in premenopausal women (RECONNECT Phase III trials)
- Octreotide for acromegaly and carcinoid syndrome
Moderate Evidence
Criteria — at least two of the following:
- One or more human RCTs with positive results, but limited in number, sample size, or scope
- Multiple well-designed observational studies in humans with consistent findings
- Extensive and consistent preclinical data supporting a plausible mechanism
- At least Phase II clinical trials completed or ongoing
- Data from more than one independent research group
What this means for readers: There is meaningful evidence suggesting this peptide may be effective for this indication, but the evidence is not yet definitive. Further human trials are needed.
Examples:
- BPC-157 for tissue repair (extensive animal data from multiple models, very limited formal human trial data, but plausible mechanism)
- Thymosin beta-4 for corneal wound healing (Phase II trials completed)
- GHK-Cu for skin regeneration (multiple in vivo and some human studies)
Preliminary Evidence
Criteria — at least one of the following:
- Early-stage human studies (Phase I, pilot studies, or small case series)
- Consistent positive results from animal studies across multiple independent research groups
- Strong mechanistic rationale supported by in vitro data, but limited in vivo confirmation
- Research primarily from a single geographic region or research group, but with plausible methodology
What this means for readers: There are signals that this peptide may have the claimed effect, but the evidence is far from conclusive. Animal results may not translate to humans, and early human data is too limited for confidence.
Examples:
- Semax for cognitive enhancement (Russian clinical studies, limited Western replication)
- LL-37 for wound healing (Phase I/II data, extensive in vitro)
- Ipamorelin for GH secretion (Phase II data, limited Phase III)
Preclinical Only
Criteria:
- Evidence limited to animal studies (in vivo) and/or laboratory studies (in vitro)
- No published human clinical trial data for the specific indication
- Mechanism is proposed based on laboratory findings, but human translation is uncertain
- May have data from only one research group or one type of animal model
What this means for readers: This peptide has shown promise in the lab, but there is no direct evidence it works in humans for this indication. The leap from cell culture or rodent model to human therapy is enormous, and most compounds that look promising at this stage never become successful drugs.
Examples:
- Epithalon for telomerase activation and anti-aging (Khavinson lab data primarily)
- DSIP for sleep regulation (animal models, limited and inconsistent human data)
- Many novel or recently discovered peptides
Insufficient Data
Too little published research exists to make any meaningful assessment. The peptide may be very new, or the available research may be too sparse, poorly designed, or contradictory to draw any conclusions.
Study Design Types We Evaluate
When reviewing the literature for a peptide, we consider the following study designs and weigh them accordingly:
Interventional studies (experiments):
- Randomized controlled trials (parallel group, crossover, factorial)
- Non-randomized controlled trials
- Single-arm interventional studies (no control group)
Observational studies:
- Prospective cohort studies
- Retrospective cohort studies
- Case-control studies
- Cross-sectional studies
- Ecological studies
Descriptive studies:
- Case series
- Case reports
Secondary research:
- Systematic reviews (with or without meta-analysis)
- Narrative reviews
- Umbrella reviews (reviews of systematic reviews)
Preclinical research:
- In vivo animal studies (rodent, primate, other)
- Ex vivo tissue studies
- In vitro cell culture studies
- In silico computational studies
Statistical Concepts We Consider
P-Values
The p-value represents the probability of observing results at least as extreme as those obtained, assuming the null hypothesis (no effect) is true. A p-value of 0.05 means there is a 5% chance of seeing such results if the treatment truly has no effect.
What p-values do NOT tell you:
- The probability that the hypothesis is true or false
- The size or clinical importance of the effect
- Whether the result will replicate
We flag studies that rely on borderline p-values (0.04–0.05) without adequate sample sizes, and studies that test many outcomes without adjusting for multiple comparisons.
Confidence Intervals (CIs)
A 95% confidence interval provides a range within which the true effect is likely to fall. Narrower intervals indicate more precise estimates. We prefer studies that report confidence intervals over those that only report p-values, because CIs convey both the magnitude and precision of the effect.
Key interpretation rules:
- If the 95% CI for a difference crosses zero (or for a ratio, crosses 1.0), the result is not statistically significant at the 0.05 level
- Wide CIs indicate imprecise estimates, often due to small sample sizes
- Two studies can both be "significant" but have very different effect magnitudes
Number Needed to Treat (NNT)
The number of patients who need to be treated for one additional patient to benefit. Lower NNTs indicate more effective treatments. An NNT of 1 would mean every patient benefits; an NNT of 100 means you need to treat 100 patients for one to benefit. We use NNT to contextualize clinical significance when available.
Hazard Ratios (HRs)
Used primarily in survival analyses and time-to-event studies. A hazard ratio of 0.5 means the treatment group has half the rate of the event (e.g., death, disease progression) compared to control at any given time point. An HR of 1.0 means no difference.
Absolute vs. Relative Risk Reduction
We pay close attention to whether studies report absolute or relative risk reductions, because relative measures can be misleading. If a treatment reduces risk from 2% to 1%, the relative risk reduction is 50% (sounds impressive) but the absolute risk reduction is only 1% (less impressive). Both numbers are technically correct, but the absolute reduction better conveys the practical impact.
Types of Bias We Assess
Selection Bias
Systematic differences between the groups being compared. In RCTs, this is minimized by proper randomization and allocation concealment. In observational studies, this is a major concern.
Publication Bias
Studies with positive results are more likely to be published than negative ones. This creates a distorted literature where treatments appear more effective than they truly are. We look for signs of publication bias including asymmetric funnel plots in meta-analyses and registered trials that never published results.
Funding Bias
Industry-funded studies tend to report favorable results more often than independently funded research. We note funding sources and evaluate whether the study design could have been influenced by commercial interests.
Observer/Detection Bias
Differences in how outcomes are measured or assessed between groups. Minimized by blinding assessors to group assignment.
Attrition Bias
Systematic differences in dropout rates between groups. If more patients drop out of the treatment group (due to side effects, for example), the remaining participants may not be representative.
Reporting Bias
Selective reporting of outcomes. A study may measure 15 endpoints but only report the 3 that showed significant results. We check study registrations (ClinicalTrials.gov) against published results to identify potential reporting bias.
Special Considerations: Khavinson Peptides
Several peptides discussed on PeptideInsight — including Epithalon, Vilon, Thymalin, and other short peptides — originate primarily from the laboratory of Professor Vladimir Khavinson at the Saint Petersburg Institute of Bioregulation and Gerontology. We apply additional scrutiny to these compounds because:
- The majority of published research comes from a single research group
- Many studies are published in Russian-language journals with limited international peer review
- Some claimed mechanisms (e.g., direct DNA interaction by tetrapeptides) are not well-supported by independent biochemical research
- The peptides are marketed commercially in Russia, creating potential conflicts of interest
- Independent replication by Western research groups is very limited
This does not mean these peptides are ineffective, but it means the evidence should be interpreted with greater caution. We explicitly note this single-source limitation in our evidence assessments.
Regulatory Status Categories
We also track the regulatory status of each peptide:
- FDA-approved: The peptide has been approved by the U.S. Food and Drug Administration for at least one specific indication
- EMA-approved: Approved by the European Medicines Agency
- Approved in other jurisdictions: Approved in specific countries (e.g., Russia, Japan, Australia) but not FDA/EMA
- In clinical trials: Currently being tested in registered human clinical trials (Phase I, II, or III)
- Investigational: Being actively researched but not yet in formal clinical trials
- Research compound: Available for laboratory research; not approved for human use
- Discontinued: Development halted due to safety concerns, lack of efficacy, or commercial reasons
How We Assign and Update Grades
Our grading process follows these steps:
- Literature search: We search PubMed, Google Scholar, ClinicalTrials.gov, Cochrane Library, and EMBASE for all published research on each peptide-indication pair
- Study cataloging: We catalog each relevant study by design, sample size, quality, and funding source
- Quality assessment: We evaluate study quality using established frameworks (Cochrane Risk of Bias tool for RCTs, Newcastle-Ottawa Scale for observational studies)
- Replication check: We assess whether findings have been replicated by independent research groups in different settings
- Consistency evaluation: We evaluate whether the totality of evidence points in the same direction or is contradictory
- Grade assignment: We assign a grade based on the totality of available evidence, weighting higher-quality evidence more heavily
- Peer review: Grades are reviewed by at least one additional team member before publication
- Periodic updates: We review and update grades as significant new research is published, with the date of last review shown on each peptide page
Limitations of Our System
- Evidence grades reflect the current state of published research and may change as new studies emerge
- A low evidence grade does not mean a peptide is ineffective — it means insufficient research exists to draw firm conclusions
- A high evidence grade for one application does not extend to other claimed uses of the same peptide
- Our system cannot fully account for unpublished data, ongoing trials not yet reported, or research in languages we cannot access
- We aim for objectivity, but all grading systems involve some degree of judgment
- Our grading is not a substitute for clinical decision-making by qualified healthcare providers