The RAPPIDS Residential Price Index
Construction, estimation and interpretation of a repeat-sales price index for Australian residential property.
- Effective
- August 2026
- Data as at
- 2 August 2026
- Estimator
- Weighted BMN
- Base period
- January 2010 = 100
1Purpose and scope
RAPPIDS — the Residential Asset Price Paired Index Data Series — measures the change in the price of Australian residential property over time, holding the properties themselves constant. It answers one question: what would a dwelling worth 100 in January 2010 be worth now?
It is a price index, not a valuation model. It does not estimate what any particular property is worth, and it should not be used as one. Its unit of observation is a pair of sales of the same dwelling, and its output is a single number per market per month with an explicit uncertainty attached to it.
This paper states how that number is produced, what is excluded from it and why, how precise it is, and where it is known to be wrong. It is written so that a reader with the same source data could reproduce the series.
2Why repeat sales
The most commonly quoted housing measure is a median sale price. Its defect is that it moves when the mix of what sells changes, with no property changing in value. A quarter in which more apartments than usual transact will show a falling median in a rising market. A quarter in which the top of the market goes quiet will show the same. The median is a statistic about sellers as much as about prices.
A repeat-sales index removes the problem by construction. It only ever compares a property against itself: the observation is not “a house in Manly sold for $3.4m” but “this house sold for $2.1m in 2016 and $3.4m in 2024”. Location, land size, aspect, orientation, street and every other fixed attribute cancel out of the comparison. What remains is movement in the price of housing.
The method is the same one underlying the S&P Case–Shiller indices in the United States and the FHFA House Price Index, and it originates with Bailey, Muth and Nourse (1963). The cost of the approach is that it discards every property that has sold only once, which is most of them; the benefit is that what survives is directly comparable across time.
3The index family
For each covered market the series is published monthly at four levels of aggregation: all dwellings, houses, apartments and townhouses. Each is estimated independently from its own pair set rather than derived by weighting the others, so the house and apartment series are genuinely separate measurements and may diverge.
| Attribute | Specification |
|---|---|
| Frequency | Monthly |
| Base period | January 2010 = 100 |
| Estimator | Bailey–Muth–Nourse regression, interval-weighted |
| Segmentation | Region × dwelling type (all, house, apartment, townhouse) |
| Date convention | Contract date |
| Published with | Standard error of the index level |
| Revision | Every period revises on each recomputation (§8) |
Current coverage of the primary series, Sydney all dwellings, as published:
| Measure | Value |
|---|---|
| First period estimated | January 1990 |
| Last period published | August 2026 |
| Periods in series | 441 |
| Index level, last published period | 246.6 (±14.2 at 95%) |
| Pairs supporting that period | 4,297 |
Markets outside Sydney are held to the same publication threshold (§8) and most do not yet clear it. Where a market fails the threshold, no chart is drawn and no figure is quoted; the site states that the threshold is unmet.
4Source data
The index is built from publicly available records of residential sales in Australia: the sale price, the sale date, the dwelling type and the property’s prior transaction history. No confidential, licensed or subscriber data is used, and no personal information about buyers or sellers enters the index.
Dates are contract dates — the date the sale was agreed — not settlement dates. This is a substantive difference from settlement-based measures, which include those built on land-titles data. Standard residential settlement in New South Wales is 42 days, so a contract-dated index observes a turn in the market roughly six weeks before a settlement-dated one records it. §10 quantifies what that looks like in practice.
Prices are recorded as transacted, in nominal dollars, with no adjustment for inflation, incentives or vendor contributions.
Not every sale is advertised with its price. Where an agent does not disclose one, the figure is established from public records and is exact to within $1,000 — a tolerance four orders of magnitude below a Sydney median, and immaterial to any return computed from it. Prices obtained this way are marked wherever they appear, because whether a sale was advertised with its price is itself informative: among listings carrying a return, undisclosed sales are loss-making at close to twice the rate of disclosed ones. No such price is excluded from the index, and none is treated differently in estimation.
5Pair construction and eligibility
A property’s transaction history is sorted into chronological order and split into consecutive legs, each leg being one sale to the next. A property with four recorded sales contributes three pairs, not six — non-adjacent combinations would reuse the same price movements and overstate the effective sample.
A pair enters the estimation only if it satisfies all of:
- Both prices are genuineBoth legs carry a recorded sale price above zero. Records carrying a placeholder rather than a price are dropped, and a current sale at or below $100,000 is treated as absent. The artefacts here are not small numbers but round ones — 68 Sydney records sat at exactly $10,000, $50,000 or $100,000, every one of them a substantial property — and a $10,000 sale after a $2,820,000 one annualises to −33.8% over a long hold, passing the return bound below. This magnitude floor is deliberately not applied to historical records: a $10,150 sale in 1971 is real, and long-horizon pairs of that kind anchor the early index.
- The interval is positiveThe two sales fall in different months. A property bought and resold within one month carries no information about movement between periods and is dropped rather than contributing a zero-length row.
- The first sale is 1990 or laterEarlier records exist but are too sparse to identify monthly periods, and their reliability is materially lower.
- The implied return is physically plausibleThe pair’s compound annual return must fall within −35% to +50%, widened to −60% to +100% for holds under twelve months where genuine short-run swings annualise to large numbers.
The return bound is the only rule that judges whether a recorded movement is real. A repeat-sales index assumes the asset did not change between its two sales. Knock-down rebuilds and vacant-land-then-completed-house resales violate that assumption in one direction: they look like enormous gains. Greenfield land-then-house pairs routinely imply 140–180% per annum. Left in, they would bias the index upward without limit, and the upper bound removes them.
The bounds do not remove subdivisions, which run the other way and are modest in size. When a house is split and one resulting lot resells, the recorded fall is a fraction, not a multiple. 7 Sunnyside Avenue, Lilyfield sold for $2,290,000 in March 2024 and one of its lots for $1,550,000 in July 2026 — −15.4% per annum, which passes the −35% bound and sits in the regression. §5.1 describes what is being done about it.
The bounds are wide and close to symmetric. A narrow or asymmetric bound would discard observations from a genuine downturn, and a filter that removes falls but keeps rises produces an index that cannot show a fall. Sydney dwelling values fell at roughly 15% per annum through parts of 2018; these bounds retain that and still remove the physically impossible.
A second, related case is the building-level transaction: a whole block or development site sells, and the sale is recorded against every individual unit in it. One observed example put a $32.3m sale into the history of four separate apartments, each of which then appeared to have “lost” some $31.7m. These are caught by the same plausibility rule, and the affected properties are shown across the site without a return rather than with a false one.
On the Sydney sample as at 2 August 2026, 27,133 candidate legs were constructed and 26,094 retained — 3.8% removed by the rules above.
5.1 Subdivisions, and what the bounds miss
A subdivision is the hardest case this index faces, because the source data actively misreports it: the pre-subdivision sale is attached to the post-subdivision parcel, so one record shows a property that appears to have fallen sharply in value. Nothing in the transaction history marks it. There are no per-sale attributes, the sale type is the literal string “Sale” on 41,999 of 42,000 records, and the land area available is a single current figure rather than one per sale.
The one artefact that does record it is the floor plan attached to the earlier sale, where agents often print a land size. A local vision model reads that figure and it is compared against the parcel recorded today; a fall beyond 20% is treated as a subdivision. On the Lilyfield property above the earlier plan prints 350m² against 153m² now, a 56% loss of land, while the dwelling itself is barely changed — two bedrooms then, and the same rooms. Any test counting rooms sees nothing.
Two limits on this, both material. Only a minority of plans print a land size at all — of 21 plans read in the first pass, one produced a usable land comparison, six were apartment plans where land is not the relevant question, and the rest printed no figure. And coverage depends on a plan existing for the earlier sale, which holds for roughly a quarter of properties with a return.
| Holding period | Pairs | Share |
|---|---|---|
| Under 2 years | 1,970 | 7.5% |
| 2 to 5 years | 6,380 | 24.4% |
| 5 to 10 years | 8,100 | 31.0% |
| 10 to 20 years | 6,812 | 26.1% |
| 20 years or more | 2,857 | 10.9% |
| Median hold | 7 years 7 months |
6Estimation
For a pair of sales of the same property at periods t₀ < t₁, with a log price index β:
This is a linear model. Each observation contributes one row to a design matrix holding −1 in the buy period’s column, +1 in the sell period’s, and zero everywhere else. Estimating it by least squares recovers β as a log price index, and the published level is:
The decisive property is the row of zeros. The regression assumes nothing about how a pair’s return distributes across the periods it spans; a 1996–2026 pair constrains the difference between those two months and is silent about the 359 months in between. Methods that instead spread a pair’s average return evenly over its span cannot recover time variation at all — they produce a smooth exponential by construction, and will show no downturn however severe the downturn was.
Only differences of β are identified, so the first period is pinned to zero and the level restored at the rebasing step. Since each row has exactly two non-zero entries, the normal equations XᵀX and Xᵀy are accumulated directly in time linear in the number of pairs, without materialising the design matrix — which for Sydney would be 26,119 × 440.
6.1 Smoothing
Pairs are not evenly distributed through time; they cluster in recent years, and a month in the early 1990s may be identified by very few observations. Unconstrained least squares lets those periods wander, producing month-to-month zig-zags that are sampling noise rather than market movement.
A ridge penalty is therefore applied to the second differences of β — that is, to the curvature of the index, not to its level or its slope. Penalising curvature pulls thinly-evidenced stretches toward the local trend while leaving a genuine turning point free to turn, because a sustained change of direction supported by data will pay the penalty and still improve fit. A linear trend is unpenalised entirely.
The smoothing weight is the one discretionary parameter in the estimator, so its influence is stated rather than assumed. Measured on the 2017–19 Sydney drawdown:
| Smoothing weight | Peak | Trough | Drawdown |
|---|---|---|---|
| λ = 0 (textbook BMN) | May 2018 | June 2019 | −12.7% |
| λ = 8 (published) | August 2017 | June 2019 | −11.7% |
| λ = 100 (heavy) | August 2017 | June 2019 | −10.5% |
A twelve-fold increase in the penalty moves the measured drawdown by 1.2 percentage points and does not move the trough at all. The published setting is not doing the work of producing the result.
6.2 Interval weighting
A house held for twenty years drifts further from the market average than one held for two: idiosyncratic factors — a renovation, a deteriorating roof, a new arterial road — accumulate with time. Residual variance therefore grows with the holding interval, which violates the equal-variance assumption of ordinary least squares and lets long-hold pairs exert more influence than their information content warrants.
The correction is the second and third stages of the Case–Shiller procedure. Squared residuals from the first fit are regressed on the holding interval to estimate how variance scales with time, and the index is re-estimated weighting each pair by the inverse of its fitted variance. Short-hold pairs, which carry the most reliable information about the periods they span, are weighted up accordingly.
7Precision and confidence bands
Every published period carries a standard error of the index level, derived from the regression’s covariance matrix and converted from log scale to level scale by the delta method. A 95% interval is the level plus or minus 1.96 standard errors.
These bands are wide. A repeat-sales index estimated from tens of thousands of pairs across four decades carries substantial uncertainty, and publishing a bare point estimate conceals that rather than eliminating it.
The standard error is not zero at the base period. Identification pins the first period of the series to zero and every later standard error is chained from there; rebasing afterwards moves the levels but not that anchor. On the Sydney series the base period carries a standard error of 10.11.
Two things follow. Uncertainty grows with distance from the start of the series rather than from the base, so the band is narrowest in 1990 and widest today — the opposite of what a chart labelled “2010 = 100” suggests. And extending the series further back inflates every band, because each additional period is one more estimated increment between the anchor and the present: admitting pairs from 1980 widens the 2010 band from ±19.8 to ±72.7 while moving the index itself by nothing.
This is a further reason to read growth rates — a ratio of two levels, in which common error partly cancels — in preference to levels.
8Publication and revision
A repeat-sales index is always thinnest at the end of the sample. A pair only spans the months between its two sales, and nearly every pair ends in the current year, so recent months rest on far fewer observations than older ones. Sydney runs some 12,000 pairs through January and around 700 by July of the same year. Read without that context, the tail-off looks like a market decline when it is mostly a shrinking sample.
A period is therefore published only once at least 2,000 pairs span it. Periods below the threshold are estimated and stored but marked provisional, and are suppressed or visibly flagged wherever they appear. Case–Shiller publishes on a two-month lag for the same reason; this rule is the data-driven equivalent, and it produces a lag that varies with actual coverage rather than a fixed one.
A market in which no period clears the threshold is not published at all. This is not a placeholder state that will be filled with a weaker figure later — it is the correct output for a sample that cannot support a number.
Revision policy
The index is re-estimated in full whenever the underlying data changes. Because every pair contributes to the fit jointly, a newly observed sale in 2026 revises the estimate for every period its pair spans, including periods years earlier. Revisions of this kind are intrinsic to repeat-sales estimation and are not corrections of errors; they are the arrival of new evidence about the past. Revisions are largest in the most recent periods and diminish as coverage deepens.
9Rebasing
The base period is a display convention, not an estimation choice. The regression recovers relative levels; fixing a base to 100 is a division applied afterwards. Any period may therefore be adopted as the base at any time without re-estimating anything, and doing so changes no growth rate, no turning point and no confidence interval other than by the same constant factor.
The published base is January 2010 = 100.
10Validation
An index that cannot be checked is an assertion. Three checks are run against evidence the index does not itself use.
10.1 Recovery of known turning points
The series independently reproduces the documented cycles of the Sydney market: the 2017–19 downturn, at −11.7% peak to trough; the 2020 pandemic dip; the 2021 boom, at +16.9% over the year to May 2021; and the 2022 rate-driven correction, at −11.6% from October 2021 to November 2022. None of these are inputs. An estimator that failed to show them — as the flat-spread predecessor to this one did — is detectably broken regardless of how smooth its output looks.
| 12 months to May | RAPPIDS, all dwellings |
|---|---|
| 2019 | −9.9% |
| 2020 | +4.4% |
| 2021 | +16.9% |
| 2022 | +5.4% |
| 2023 | −1.4% |
| 2024 | +1.6% |
| 2025 | +9.0% |
| 2026 | +7.2% |
10.2 Comparison against an independent benchmark
Compared against a major commercial settlement-dated home value index over the 2017–19 cycle, RAPPIDS turns one to two months earlier and records a shallower fall: −11.7% against −14.4% peak to trough. Both differences are expected and both have identified causes. The timing difference is the contract-versus-settlement convention of §4, and the estimated lead shrinks as smoothing is removed — from about three months on a twelve-month average to one month on raw monthly levels — which is what a genuine dating difference, rather than an artefact, should do. The magnitude difference is consistent with the quality drift described in §11.
10.3 Coincident indicators
Monthly movements are checked for consistency with auction clearance rates over the same window, an indicator with no methodological relationship to repeat-sales estimation.
11Known limitations
Every one of these is a real defect. They are listed because a reader cannot calibrate a number without them.
- Quality driftThere are no hedonic controls. A property renovated between its two sales attributes the improvement to the market, and properties that transact twice skew toward improved stock. Measured against an independent benchmark, this appears to run roughly +0.5% per annum high. This is the largest known bias in the index and it is directional, not random.
- Transaction-frequency biasOnly properties that have sold at least twice are observable. Investor stock and short-hold properties are over-represented; long-held owner-occupied family homes are under-represented, and any respect in which they behave differently is invisible to the index.
- Improvement filtering is blunt, and asymmetricThe plausibility bounds of §5 are a crude proxy for detecting a changed asset. They catch the direction that produces enormous gains and largely miss the one that produces modest falls: a subdivision reads as −15%, not −90%, and passes. §5.1 covers the supplementary check for that case and its coverage limits. Case–Shiller’s explicit renovation flagging is more principled; this is not that.
- Monthly resolution exceeds the dataMonth-to-month movement in this index is substantially noise. Annual growth is where the signal is. Quarterly or annual changes, or a short trailing average, should be preferred over raw monthly levels — and the confidence band makes the reason visible.
- Sample composition is not weighted to the marketThe pair set is not stratified or reweighted to match the composition of the dwelling stock. A market whose repeat-sale properties differ systematically from its housing stock will be measured with that difference embedded.
- Coverage is unevenSydney is the only market with sufficient depth for the full series. Others are published only where they clear the threshold in §8, and most do not.
12Glossary
- Repeat-sale pair
- Two consecutive recorded sales of the same property, treated as one observation of price movement between those two dates.
- BMN
- Bailey–Muth–Nourse: the regression method underlying repeat-sales indices, introduced in 1963.
- Contract date
- The date a sale was agreed, as distinct from settlement, when title transfers. This index is dated on contract.
- Base period
- The period set to 100. A display convention with no effect on estimation (§9).
- Standard error
- The estimated sampling uncertainty of an index level. A 95% interval is ±1.96 standard errors.
- Provisional period
- A period estimated but not published because fewer than the required number of pairs span it (§8).
- Pair count
- The number of retained pairs spanning a period — the observations identifying that period’s estimate.
- CAGR
- Compound annual growth rate: the constant annual rate that connects two prices over the holding period between them.
- Hedonic control
- Adjustment for measured property attributes. This index uses none; the repeat-sales design controls for fixed attributes implicitly instead.
13References
- [1]Bailey, M.J., Muth, R.F. & Nourse, H.O. (1963). “A Regression Method for Real Estate Price Index Construction.” Journal of the American Statistical Association 58(304), 933–942. The original method. The estimator in §6 is this paper. Link
- [2]Case, K.E. & Shiller, R.J. (1987). “Prices of Single Family Homes Since 1970: New Indexes for Four Cities.” New England Economic Review, September/October, 45–56. Introduced the interval weighting applied in §6.2.
- [3]Case, K.E. & Shiller, R.J. (1989). “The Efficiency of the Market for Single-Family Homes.” American Economic Review 79(1), 125–137.
- [4]Wharton School, University of Pennsylvania. Repeat Sales House Price Index Methodology. The clearest derivation of the design matrix. Link
- [5]S&P Dow Jones Indices. S&P CoreLogic Case-Shiller Home Price Indices Methodology. The practitioner specification, including three-stage weighting. Link
- [6]Federal Reserve Bank of Philadelphia (2014). House Price Indexes: Methodology and Revisions. Plain-English account of why simpler approaches fail, and of revision behaviour (§8). Link
Version 1.2, effective August 2026. Coverage figures are as at 2 August 2026 unless read live from the published series.
The index is an estimate published with its uncertainty. It is not a valuation of any individual property and is not financial advice.