RESEARCH PEPTIDE FUNDAMENTALS / MATRIX
Three Peptides Compared on How They Were Studied
The same six appraisal dimensions applied to retatrutide, BPC-157 and thymosin alpha-1 — which produces a very different ranking from comparing their mechanisms.
The short version
Most comparisons of research peptides line up what each one does. This one lines up how well each one has been tested, which is a different question and often a more useful one. The six things being compared are simple. Was there a comparison group? Did anyone know who was getting what? How many people took part? Was the thing measured something that matters directly, like whether people lived, or a stand-in like a blood test? Has anyone repeated it? And has the final, confirming study been done? On those measures the ranking flips. Retatrutide has excellent studies that have not finished. Thymosin alpha-1 has the single best-built study on this site, and it came back negative. BPC-157 has the largest pile of research and the weakest human evidence in the group. Quality of testing and volume of literature are not the same thing.
The evidence-quality matrix
| Appraisal dimension | Retatrutide | BPC-157 | Thymosin Alpha-1 |
|---|---|---|---|
| Strongest published human design | Randomised, double-blind, placebo-controlled phase 2; one trial also active-controlled [4][5] | Uncontrolled first-in-human safety pilot [8] | Multicentre, double-blinded, randomised, placebo-controlled phase 3 [13] |
| Largest human sample | 338 adults [4] | 2 adults [8] | 1,106 adults [13] |
| Blinding | Double-blind throughout [4][5][6] | None | Double-blind [13]; the earlier trial was single-blind [17] |
| Control arm | Placebo, plus an active comparator in type 2 diabetes [5] | None [8] | Placebo [13] |
| Primary endpoint type | Surrogate markers: body weight, HbA1c, liver fat [3][4][5] | Safety biomarkers only [8] | A hard clinical event: 28-day all-cause mortality [13] |
| Dose-response demonstrated | Yes, ordered across four ascending doses [3] | In rodents only [12] | Not the design question in the pivotal trial |
| Confirmatory trial status | Phase 3 ongoing [1] | Never conducted [9] | Completed, and null [13] |
| Dominant evidence species | Human [3][4][5][6][7] | Rat, dog, chick membrane, cultured cells [10][11][12] | Human [13][15][17] |
| Independent replication | Industry-run programme; confirmation pending [1] | Foundational literature concentrated in one research group [9] | Replicated at higher rigour; the signal did not hold [13][17] |
| The most useful question to ask next | Does the marker predict the outcome? | Has anyone unrelated repeated this in a person? | Which of the two trials is being quoted? |

Control and blinding
A control arm answers the question "compared with what?", and without one, almost nothing observed can be attributed to the compound. Retatrutide's trials all carry a placebo arm, and the phase 2 type 2 diabetes trial carries an active comparator as well [5] — a stronger construction, because it tests not whether the drug does something but whether it does more than an existing option. Thymosin alpha-1's pivotal trial is placebo-controlled [13]. The BPC-157 human record has no control arm anywhere in it [8].
Blinding then determines how much the control arm is worth. The comparison available here is unusually direct, because thymosin alpha-1 was tested twice in the same condition on the same primary endpoint at two levels of blinding. The single-blind trial, with 361 patients, reported 26.0 percent 28-day mortality against 35.0 percent in controls [17]. The double-blind trial, with 1,106 patients, reported 23.4 percent against 24.1 percent [13]. Both were randomised and multicentre. The most economical explanation for the difference is not that the peptide stopped working; it is that unblinded clinical judgement in a critical-care setting is a powerful source of differential care between arms, and that removing it removed the effect.
Sample size, and what a number can detect
The human sample sizes across this site span nearly three orders of magnitude: 2 [8], 72 [6], 98 [3], 281 [5], 338 [4] and 1,106 [13]. Sample size is not a measure of prestige. It is a measure of what a study is capable of noticing, and reading it that way makes several papers easier to place.
A two-person report [8] is capable of noticing effects so common they would appear in nearly everyone exposed and nothing else — which is precisely the job of a first-in-human safety pilot and a reasonable design for that purpose. A 72-person phase 1b trial [6] is sized to characterise pharmacokinetics and tolerability, which is why its efficacy figure arrives with a confidence interval wide enough to make the point itself. A trial of several hundred [4][5] can detect a large efficacy difference on a marker that moves reliably. A trial of a thousand-plus [13] is sized to detect a plausible difference in a rare-ish binary event such as death within 28 days, which requires far more participants than a continuous marker does.
The corollary matters more than the rule: a study that fails to find an effect it was never large enough to detect has not shown the effect is absent. The distinction between an underpowered study and a genuine null is entirely a matter of design, not of result.
Endpoints: a marker or the thing itself
This is where the three compounds separate most sharply, and where the ranking most surprises a reader coming from the pharmacology.
Retatrutide's most impressive figures are all surrogates. A relative liver-fat reduction of 82.4 percent with 86 percent of participants reaching normal liver fat [3], a body-weight change of -24.2 percent [4], an HbA1c reduction of 2.02 percent [5] — every one of these is a measurement chosen because it moves quickly, quantifies precisely and can be captured inside a trial of practical length. Surrogates are legitimate and necessary; they are also a bet that moving the marker moves the outcome, and that bet has to be settled by a separate trial powered for events. Retatrutide's has not reported.
Thymosin alpha-1's pivotal trial measured death [13]. There is no interpretive gap between the endpoint and what anyone cares about, which is exactly why the null carries the weight it does. A negative result on a hard endpoint in a well-powered blinded trial is among the most informative things a literature can contain, and it is systematically under-cited relative to positive results on markers.
BPC-157's human record measures neither. The pilot reported cardiac, hepatic, renal, thyroid and glucose biomarkers and observed no changes [8] — safety parameters, not efficacy endpoints of any kind. The rodent work does report outcomes, such as an ulcer-formation inhibition ratio of 45.7 to 65.6 percent [12], but those are outcomes in rats.
Replication, independence and who ran the study
The three compounds illustrate three distinct provenance situations, and they are not equivalent.
Retatrutide's trials are a sponsor-run development programme [1][4][5][6]. The designs are strong and the reporting is detailed, and the appropriate response is not suspicion but specificity: read the pre-registered protocol, check whether the reported primary endpoint is the registered one, and note that confirmation is still pending.
BPC-157's literature is concentrated in a single research group and its collaborators, which newer reviewers flag explicitly [9]. This is a structural weakness rather than an allegation. Independent replication exists to catch systematic errors that everyone inside one programme shares — an assay quirk, a model artefact, an analytic habit — and a literature that has never left its originating group has never been exposed to that check.
Thymosin alpha-1 has the outcome that replication exists to produce. An earlier signal was tested by a larger, better-blinded, independently conducted trial and was not confirmed [13][17]. That is not a failure of the scientific process; it is the process working, and it is the reason the earlier positive trials and the meta-analyses built on them can now be read at their correct weight rather than at their apparent one.