← Neighborhood Guides

How We Measure Safety Without Grading Neighborhoods

Most sites either dropped crime data or reduced it to a letter grade. We built our safety signal around the research finding both approaches ignore: crime concentrates on a tiny fraction of blocks.

Ask renters what they want to know about a place before signing a lease and safety lands at or near the top of every survey. Then try to find it on the sites where people actually search. You mostly can't, and the reason why is one of the more interesting unresolved arguments in real estate.

In December 2021, Redfin and Realtor.com both announced they would not show neighborhood crime data, and Trulia removed its crime maps days later. These weren't legal-compliance footnotes; Redfin published its reasoning. The available data couldn't accurately answer the question renters were asking, and presenting it anyway risked reinforcing the racial bias already documented in how people perceive safety. The industry's other answer, the letter grade that stamps a C-minus on an entire zip code, is arguably worse: it takes everything wrong with a neighborhood average and gives it the visual authority of a report card.

So the field split into two camps: show nothing, or show something blunt and unfair. We think both camps are responding to a real problem and both are missing the same fact, one of the most consistently replicated findings in criminology: crime does not spread evenly across neighborhoods. It concentrates, block by block, on a tiny fraction of places. A method built around that fact can say something useful without grading anyone's neighborhood. This article walks through ours, end to end, using our real data.

Why good intentions aren't enough

Start with why a naive crime map is dangerous even when nobody involved means harm. The core problem is that perceived safety and measured incidents are different quantities, and the gap between them is not random.

In a landmark Chicago study, sociologists Robert Sampson and Stephen Raudenbush measured physical disorder directly, sending trained observers with video down more than 20,000 block faces, then compared what was objectively there with what residents said they saw. Observed disorder mattered, but the racial and economic composition of a neighborhood predicted perceived disorder more strongly than what was actually on the street. The effect held for respondents of every race.

A second study makes the point even more directly. Lincoln Quillian and Devah Pager compared residents' perceptions of their neighborhood's crime level against police statistics in Chicago, Seattle, and Baltimore. The percentage of young Black men living nearby predicted how much crime people believed there was, even after controlling for the crime that police data said was actually occurring.

This is the research context the portals were reacting to, and it has a sharp implication for anyone building a safety feature: a product that asks users to trust their gut, or that packages a whole neighborhood into one grade, isn't neutral. It hands a megaphone to exactly the perceptual bias the research documents. Whatever we built had to start from reported incidents, not reputation, and had to resolve finer than a neighborhood.

What reported crime actually measures

Reported incidents are a real signal. Robbery and burglary reports correspond to events that harmed someone, and their geography is far from random. But the numbers also carry history. Places with elevated reported crime today are overwhelmingly places that experienced decades of concentrated poverty and disinvestment, much of it produced by explicit policy. Policing intensity feeds the data too: more patrols in an area generate more recorded incidents for some offense types, which can justify more patrols. None of that is a reason to hide the data. It is a reason to be careful about which offenses you count and what you claim the numbers mean.

Two design rules follow directly. First, we count only five offense categories: homicide, robbery, aggravated assault, burglary, and motor vehicle theft. These are serious, consistently defined, and mostly victim-reported, which makes them the categories least sensitive to enforcement discretion. No drug offenses, no loitering or nuisance categories, no arrest-only records, because those track where police look at least as much as where harm happens. Second, demographic data never enters the pipeline. Race and income are not inputs, not adjustments, and not context layers on any map we draw, including the ones in this article.

The phrase doing the work there is victim-reported. A burglary or a stolen car enters the record because someone it happened to picks up the phone, so the count barely depends on how many patrol cars are nearby. A drug or loitering charge is different: an officer has to be present and choose to act, so its count climbs wherever police attention is heaviest. Try it below.

Two kinds of crime statistic

Drag to change how heavily this block is patrolled. Watch the two columns come apart.

Police presenceTypical patrol level · 1.0x
Reported by a victim
Someone calls it in
We count
100
BurglaryRobberyAssaultHomicideCar theft
Only counted if police see it
An officer must act
We drop
100
Drug possessionLoiteringNuisanceDisorderly conduct
Drag the slider. The blue column tracks what happened to people. The red column tracks where police were looking.

That is the bias in one picture. Our safety signal counts only the blue column, five serious offenses a victim reports, and drops the red one entirely, because its geography reflects patrol patterns as much as harm.

Illustrative response shapes, indexed to a typical patrol level. Victim-reported offenses are near-flat to police presence; discretion-based offenses scale with it. This is the mechanism our offense filter is built to avoid.

Crime is not a neighborhood property

Here is the empirical finding our whole approach is built around. Criminologist David Weisburd, after tracking incidents street segment by street segment across cities for decades, proposed what he called the law of crime concentration: in city after city, roughly half of all reported crime occurs on about 2 to 6 percent of street segments. The pattern replicates remarkably well, from Seattle to Tel Aviv, and it is stable over time. Most blocks in any neighborhood, including neighborhoods with elevated overall rates, see little or no serious crime. A small set of specific places accounts for most of it.

We checked this against our own data rather than taking the literature's word for it. Our unit is the hexagonal cell, about 320 meters across, which is roughly a dozen street segments, so concentration measured at hex scale should look milder than Weisburd's segment-scale numbers. It is still stark.

of hexes

7.9%

hold half of all violent-crime reports in the City of LA over the last three years. 28% of hexes recorded zero.

Scroll through what that looks like on the ground, in a three-kilometer circle around Hollywood and Highland. This is our real data: every dot below is a police report from the last three years.

Scroll to move through the map
Los Angeles

Window

3 years

Jul 2023 to Jul 2026

Reports

32,881

every offense type

Three years of police reports in one slice of Hollywood: 32,881 incidents of every type, from stolen bikes to felony assaults. Plotted raw, it reads as noise, and as a smear: the whole area looks the same. This is roughly what a naive crime map shows you.

Reported incidents around Hollywood, July 2023 to July 2026. LAPD public data; dot positions are the department's hundred-block approximations, thinned proportionally for rendering.

This is why a neighborhood grade fails in both directions at once. Average the corridor into the quiet blocks and you stigmatize thousands of residents who live on streets with effectively no serious crime. The same average understates the specific corridor a renter might actually be choosing between. The neighborhood is simply the wrong unit. So we don't use it.

The pipeline, step by step

What follows is the actual pipeline that scores every listing on Saktoo, run on the same Hollywood frame. Nothing here is a simplified illustration; the maps are the pipeline's real intermediate outputs.

Scroll to move through the map
Los Angeles

Before

32,881

all reports

After

7,752

five serious categories

Step one: the offense filter. Of 32,881 reports in this frame, 7,752 survive, about one in four. What stays is the five serious categories: homicide, robbery, aggravated assault, burglary, motor vehicle theft. What fades out is everything whose geography tracks enforcement attention as much as harm: drug offenses, nuisance categories, arrest-only records.

The real scoring pipeline on real data: offense filter, per-capita rates, empirical Bayes smoothing, pooled percentiles.

Step three, the smoothing, is the part most publishers skip and never mention, and it is worth slowing down on because it sounds like a trick and is not. The problem it solves is the one from step two: that hex posting 385 incidents per 1,000 residents got there on a tiny population, so the rate is mostly noise. Shrinkage asks a simple question, how much should we trust a number built on this few people, and answers it with the population itself. Drag the slider below to feel it.

Why a rate needs a crowd behind it

1The problem

A crime rate is just incidents divided by residents. Divide by a big population and the rate is stable. Divide by a tiny one and a single unlucky incident explodes into a huge number. This hex shows 385 per 1,000, but we cannot tell from the rate alone whether that is real danger or one event over too few people.

2How we handle it: drag the crowd

So we ask a second question: how many people actually stand behind that number? The more residents, the more we trust the hex's own rate. The fewer, the more we fall back on what its district looks like.

People living in this hex60
91
Raw rate · 385
what this hex claims
Our estimate · 91
rides between the two
District avg · 32
what neighbors look like
Trust earned by this hex
17%too few people to trust the raw rate, so we mostly follow the district
3The math: a weighted average
17%of 385this hex
+
83%of 32its district
=
91per 1,000our estimate

That trust weight is not a guess: it is exactly people / (people + 298), where 298 is the typical hex population. Small hex, small weight, so the district dominates and the wild rate gets pulled back. Big hex, the weight approaches 100% and the raw rate stands on its own. Statisticians call it empirical Bayes shrinkage; it is the same instinct as the coin-flip rule: three heads in four flips proves nothing, three hundred in four hundred does.

The estimate is the real shrinkage rule, blending the hex's raw rate with its district average by a weight of n / (n + k), where k is the median hex population. Numbers pinned to the article's loudest hex; raw rate held fixed to isolate the effect of population.

Now watch the same thing happen across a whole map. Toggle between the raw and smoothed rate below. The outlined hex, that tiny residential population next to a busy commercial strip, gets pulled back to something defensible, while the genuinely elevated corridor barely moves. That asymmetry is the whole point: shrinkage removes noise, not signal.

One honesty rule sits underneath all of this: we only claim block-level knowledge where block-level data exists. LAPD and the LA County Sheriff publish incident coordinates, so the City of LA, unincorporated LA County, and the sheriff-contract cities get the full hex pipeline, about 6.35 million residents of street-level coverage. Cities whose police departments publish only citywide totals, which includes most of Orange County, are scored from California Department of Justice agency counts instead.

That fallback is worth naming plainly, because it is in tension with everything above: for those cities we are back to the single citywide number this whole method argues against, the same coarse unit that averages a quiet block together with a busy corridor. We keep it because a clearly labeled citywide figure is more useful than either silence or a guess, but we hold it to the exact same low weight, never let it hard-filter, and phrase every sentence about those places as city-level, never block-level, so it is never mistaken for the street-level signal. Where a city self-publishes enough local detail, such as Long Beach and Santa Ana, we rescale the citywide rate across real local patterns to recover some of that lost resolution. And where we have no defensible data at all, we show nothing, because the absence of data is not evidence of zero crime.

What a renter actually sees

All of that machinery surfaces in the product as something deliberately small. The safety signal is off by default; it enters your match scores only if you turn it on in the quiz, and it never hard-filters a listing out of your results. When it is on, it contributes a small share of the overall score, never more than about eight percent even at the strongest setting, and shows up in the score breakdown as a percentile with a plain-language explanation.

The concentration flag from step four is what keeps those explanations honest. Remember the Hollywood frame: an elevated hex one block from a quiet one. If a listing sits near a localized hot spot, saying 'this area has elevated crime' would smear the whole vicinity. So the sentence changes shape. Near a concentrated hot spot, the product says:

This block has reported incident rates above most of the region, though reported incidents here cluster near the Hollywood Boulevard corridor rather than across the surrounding blocks.

Where elevated incidents really are spread across an area, it says so flatly instead:

The immediate area has reported incident rates near or somewhat above the regional average. Police-reported serious violent and property offenses here are moderately more frequent than the LA-Orange County median for the past three years.

Note what neither sentence contains: the words safe, unsafe, or dangerous, any letter grade, or any claim about who lives there. Reported incidents, a scope, a time window, a comparison. That register is a design rule, not a stylistic preference, and it applies to every surface the signal touches, including this article.

The signal's limits are part of the methodology too. Reported incidents are not all incidents; reporting rates vary by offense and by community trust in police. Resident-population denominators overstate rates in places where daytime crowds dwarf the people who sleep there, which is why a tourist strip can post a higher percentile than intuition suggests. And police data inherits the enforcement patterns of the agencies that produce it, which our offense filter reduces but cannot eliminate. We would rather state those caveats plainly than pretend a cleaner number exists.

Find out which neighborhood fits your lifestyle

Answer 10 questions. Get matched.

Take the quiz →

How we researched this

Incident data: LAPD public crime data (legacy UCR and NIBRS feeds, unioned across the department's 2024 records-system transition), LA County Sheriff Part I and II crimes via the county's open GIS layers, and California DOJ OpenJustice Crimes and Clearances for agency-level counts, all over the window July 2023 to July 2026. Populations are 2020 Census block counts. Every map in this article is the actual scoring pipeline's output on that data, baked in July 2026; incident dots use the source data's hundred-block anonymized positions and are proportionally thinned for rendering. Research citations: Weisburd's law of crime concentration and its replications (Journal of Quantitative Criminology), Sampson and Raudenbush on perceived disorder (Social Psychology Quarterly, 2004), and Quillian and Pager on race and perceived crime (American Journal of Sociology, 2001), all linked where cited. The concentration statistics quoted for the City of LA (5 percent of hexes holding 39 percent of violent-crime reports; 7.9 percent holding half; 28 percent recording zero) were computed from this same pipeline. If you spot an error in the data or the method, tell us and we will check it against the sources.

Find out which neighborhood fits your lifestyle

Answer 10 questions. Get matched.

Take the quiz →