This Corpus Was the Control
Our sister research programme at the Vatican published fifteen measurements over the course of a year, each computed from a single corpus about a single building β the standard weakness of this kind of work. This corpus is what those findings were tested against: 4,291 rated Colosseum reviews, collected in a separate effort months apart, sharing no reviews with the Vatican set. Fourteen of the fifteen held.
That matters for our own articles as much as for theirs. A finding that appears in two independent datasets about two different monuments is describing something about how people visit places, not something about one ticket office.
The Number That Came Back Identical
The Vatican team had published that longer reviews rate lower: 4.68 stars in the shortest tenth of reviews, 3.39 in the longest, a drop of 1.29. We ran the same split here and got 4.70 down to 3.41 β a drop of 1.29. The same figure to the second decimal, from a different scrape of a different monument.
Review length vs rating, measured twice in different buildings
Colosseum, shortest tenth4.70 ★
Vatican, shortest tenth4.68 ★
Colosseum, longest tenth3.41 ★
Vatican, longest tenth3.39 ★
Bars start at 3.35. Colosseum drop: 1.29 stars. Vatican drop: 1.29 stars. Neither corpus knew about the other.
4,291 and 7,714 rated on-topic items, each split into ten equal groups by character count.
Satisfied visitors write a sentence. Dissatisfied ones write paragraphs. That is not a fact about the Colosseum or the Vatican β it is a fact about review sites, and it means the most detailed review on any listing is systematically the least positive one.
The trade-off: This is the most transferable thing in either programme, and it is also the least actionable. Knowing it changes how you read reviews, not what you book.
β Do longer reviews mean worse experiences?
In both corpora measured, yes, and by the same amount. Colosseum ratings fall from 4.70 stars in the shortest tenth of reviews to 3.41 in the longest; Vatican ratings fall from 4.68 to 3.39. Both drops are 1.29 stars. When comparing tours or tickets on any platform, treat length as a signal of grievance rather than of thoroughness.
Fourteen Findings, Two Buildings
Every figure below is a distance from its own corpus average, because the two baselines differ: 4.40 here against 4.08 at the Vatican. Comparing raw star ratings between two monuments would say nothing. Comparing how far a subject pulls a review from its own norm says a great deal.
| What the review mentions | Colosseum | Vatican | Verdict |
| Guide quality | +0.18 | +0.28 | Held, closely |
| Audio guide | −0.24 | −0.10 | Held, closely |
| Visiting with children | −0.19 | −0.31 | Held, closely |
| Duration and pacing | +0.08 | +0.22 | Held, closely |
| Toilets | −0.42 | −0.62 | Held, closely |
| Engaging with the history | +0.40 | +0.62 | Held |
| Cancellations | −3.14 | −2.86 | Held |
| Price | −1.14 | −0.82 | Held |
| Meeting points | −0.95 | −0.51 | Held |
| Crowding | −0.31 | −0.60 | Held |
| Group size | −0.49 | −0.71 | Held |
| Accessibility | −0.36 | −0.80 | Held |
| Weather | −0.11 | −0.49 | Held |
| Scams and resellers | −2.45 | −1.71 | Held, worse here |
| Photography | +0.31 | −0.05 | Inverted |
The five marked as holding closely land within 0.15 of each other. Guides help by roughly the same margin in both places; audio guides underperform in both; children cost something in both; toilets drag in both. Two of our findings are notably worse here than at the Vatican β scams at β2.45 against β1.71, and meeting points at β0.95 against β0.51 β which is consistent with the Colosseum having a larger reseller ecosystem outside its gates.
The trade-off: Replication tells you a result is not an accident of one dataset. It does not tell you why the result happens, and it does not make a correlation into a cause.
The One That Flipped
Photography rates +0.31 at the Colosseum and β0.05 at the Vatican, and the explanation is architectural rather than statistical. The Sistine Chapel bans photography, so Vatican reviewers mention it in the moment they are told to put the phone away. Here, photographing the arena is most of what people came to do.
The one measurement that changed sign
Photography — Colosseum+0.31 n=573
Photography — Vatican−0.05 n=751
Same tag, opposite sign. The Sistine Chapel forbids photography, so Vatican reviewers raise it when they are being stopped. Here it is most of what people came to do.
Distance from each corpus average: 4.40 Colosseum, 4.08 Vatican.
That single row is why the exercise was worth running. Without it, a reader could take any of these measurements as describing monument visits in general. It proves that some of them describe a visitorβs relationship to a local rule, and that the honest default for an untested finding is to treat it as being about one building until shown otherwise.
β Why does photography rate positively at the Colosseum but not at the Vatican?
Because the tag is capturing different situations. Photography is prohibited inside the Sistine Chapel, so Vatican reviews raise it in the context of being stopped or of the rule being enforced, giving β0.05. At the Colosseum photography is a main activity and reviews raise it in the context of enjoying themselves, giving +0.31. Same measurement, different local rule.
What This Corpus Could Not Settle
Three of the Vatican programmeβs strongest findings β on staff, on wayfinding, and on the sensation of being herded β have no verdict here, and the reason is a limitation of our own data rather than theirs. Our TripAdvisor sample for the Colosseum venue was collected with a deliberate 1β3 star filter to surface pain points, so it had to be excluded from every comparison. Removing it took most of the monument-specific commentary with it, leaving samples too small to test.
The trade-off: That filter is what makes our pain-point articles as detailed as they are. It is also what prevents this corpus from serving as a control for anything that depends on the general tone of Colosseum reviews. Both things are true and we would rather state the cost than pretend the sample is something it is not.
β Is the Colosseum worse than the Vatican for scams and meeting points?
On these measurements, yes. Reviews mentioning scams sit 2.45 stars below the Colosseum corpus average against 1.71 below at the Vatican, and meeting points 0.95 below against 0.51. Both gaps point the same way in both corpora, so the direction is not in doubt; the larger magnitude here is consistent with a bigger reseller and street-seller presence around the monument.
The Vatican side of this comparison, written from their corpus outward, is published as
What Replicates Across Monuments, and What Does Not. Same measurements, same two datasets, opposite vantage point.
Author and Method
Research and analysis by the Intercoper Curator Team. Reviewed by Mario Dalo, founder of Intercoper.
Corpora: the Colosseum research corpus (4,291 rated on-topic items after exclusion, corpus average 4.40) and the Vatican Tour Research Corpus (22,771 items, 7,714 rated on-topic, corpus average 4.08). Assembled by separate collection runs months apart, sharing no reviews.
Exclusion applied: 1,928 TripAdvisor reviews attached to the Colosseum venue were collected with a deliberate 1β3 star filter and contain zero 4 and 5 star ratings. Every figure here excludes them. Including them would move this corpus average from 4.40 to 3.79 and make the length gradient appear twice as steep β which is the false result an earlier pass produced before the rating distribution was checked by venue.
Method: each figure is the difference between the average rating of reviews carrying an enrichment tag and the average of that corpus overall. Differences are compared rather than raw ratings, since the baselines differ by 0.32. Tags were assigned by automated per-item enrichment. Categories overlap. Sample sizes range from 49 to 3,163 and are stated in the underlying articles on each site.
Limitations: replication establishes that a result is not an artefact of one datasetβs assembly; it does not establish causation. Two corpora are not a sample of monuments. Findings on staff, wayfinding and crowd handling could not be tested here because of the exclusion above and remain single-corpus results. Trustpilot remains negatively skewed by collection design in both corpora.