.22LR Ammo Testing and Sample Sizes

grauhanen

CGN Ultra frequent flyer
Super GunNutz
Rating - 100%
178   0   0
Part 1 : Does your Rifle Really "Like" or "Dislike" the Ammo You’re Testing?

Every rimfire shooter soon understands that a group or two won’t tell him much of anything useful. That’s why he shoots a number of groups when he tests ammo.

The following, or some variation of it, happens every day at the range.:

"I did some serious lot testing today with two of my best rifles. I shot two full boxes of the same match ammo lot through each barrel in twenty 5-shot groups. After averaging the group sizes, Rifle A averaged a nice 0.350 inches, while Rifle B averaged a poorer 0..600 inches. Rifle A clearly loves this lot, but Rifle B does not."

He can honestly say “I shot the same ammo with two good rifles. One did well, the other didn’t.”

It sounds like a bulletproof, reliable experiment. The shooter didn't just fire one lucky group and call it a day; he put in the bench time, carefully shot 100 rounds in each rifle, measured every target, and calculated a mathematical average.

But despite all that hard work and experience, this shooter still did not get reliable information. It's a problem many shooters may experience.

On average, Rifle B doesn't "hate" the ammunition lot, and Rifle A doesn't have a magical preference for it. The problem usually isn't the shoooter or a lack of testing discipline -- the problem is the hidden mathematical trap built into the 5-shot group itself.
 
What’s Wrong with 5-shot groups?

Group size is the spread between the two furthest apart points of impact. It’s the group’s Extreme Spread or ES. We all use ES because it's fast and practical. But because ES only tracks the distance between the two furthest holes, it is an unstable or volatile standard of measurement (metric).

When you shoot a 5-shot group, you are only giving the ammunition five brief chances to display its round-to-round differences (noise) combined with any rifle/rest set up inconsistency that may happen from shot-to-shot (more noise). Because that sample is so small, the extreme spread of a 5-shot group swings wildly from target to target purely due to random chance.

Here is where the math traps us: because individual 5-shot groups are so unstable, even when you shoot 10 or 20 of them and average them together, the final number still carries a great deal of statistical noise.

A 100-round test made up of 5-shot groups still suffers from a huge amount of data inconsistency.


It’s a statistical coin flip. Was it chance or was it real? You can’t be sure. Coin flips are not the best way to evaluate ammo performance.
 
What should you do to be more sure?

The results you get from any single day at the range are not written in stone as a true reflection of the ammunition. If you stop testing after those two boxes, you cannot possibly know if you just witnessed a permanent mechanical truth or a temporary fluke. If you went back out the very next day and shot another two boxes per rifle, the averages could easily flip-flop completely. It’s what some may refer to as the Same Ammo, Different Results problem.

The only way to know for sure is to test further. You must increase your sample size to either confirm what you saw previously or disprove it.
 
What Causes Group Size Variation?

Why do group sizes swing so wildly from target to target? Your point of impact is never decided by just one thing. From the moment the firing pin strikes the rim, the bullet's path is decided by a combination of things:
  • Inconsistencies inside the box: Rounds may look alike but they are not identical, and in more ways than MV differences. Tiny round-to-round variations occur in how the primer compound is distributed inside the rim, propellant, casing to bullet crimp levels, and flaws with the bullet itself: differences in weight, diameter, heel asymmetry, or center of gravity variations. Such inconistencies mean trajectory differences between rounds. Lighter barrels tend to show these differences more than heavier ones because they are more sensitive.
  • The wind: Shifting air currents downrange, even ones you can't easily feel at the bench.
  • How the rifle behaves: Subtle variations in barrel vibrations, how the action sits in the stock, and minor changes in how it sits in the rest.
Some or all of these sources of variation combine together. Five shot groups that can vary in size is the result.
 
Overcoming the Problem of Wild Group Size Variation

To get a more honest, more repeatable measurement that reflects reality rather than statistical luck, we have to change the baseline unit of our testing. We need to stop using 5-shot groups for lot evaluation and switch to 10-shot groups.

A ten-shot string forces the ammunition to give its manufacturing variations twice as many opportunities to show up on a single piece of paper. It takes away the "lucky streak" factor that can make some 5-shot groups look so appealing. (Or take away the "unlucky streak" that makes some 5-shot groups look quite terrible.)

At 50 yards it’s more reliable to test ten-shot groups. Five 10-shot groups are better than ten 5-shot groups. Better still, ten 10-shot groups at 50 increases the reliability of the ES results. There’s simply less chance for random luck – good or bad – to have a big role.

Ten 10-shot groups are statistically more reliable than twenty 5-shot groups. Same number of shots but more trustworthy. Nevertheless, as hard as it is to get stable data at 50 yards using 5-shot averages, stepping back to 100 yards turns this statistical problem into a big minefield.

The next part will look at the 100-yard data problem, and why even a careful 100-round test shot in the preferred ten-shot groups can still lie to you.
 
Part 2: The 100-Yard Data Problem

In Part 1 we looked at how 5-shot groups at 50 yards can trick us into thinking a rifle somehow "loves" or "hates" a certain lot of ammunition. We saw that at 50 yards you need at least 100 total rounds in ten-shot groups before the random streaks of luck, good or bad, blend out and the truer picture is revealed.

But if you take that exact same 100-round testing and go out to 100 yards you are walking straight into a statistical trap.

At 100 yards, a standard 100-round test (two full boxes shot in ten-shot groups) may completely lie to you.
 
Why 100 Yards Changes the Rules: What works at 50 doesn't at 100

As challenging as 50 yards can be for many shooters, doubling the distance is actually more than twice the challenge. With rimfire, everything gets worse, gets more difficult, gets more magnified as distance increases.

When you double the distance, the extreme spread (group size) doesn't just double – it regularly triples. Multiple factors with .22LR combine together to make results grow in a non-linear way, becoming worse as distance grows.

1. Launch Angle Variation: a major factor that’s unseen except on the target.
  • We know that a rifle barrel vibrates like a whipping tuning fork during a shot. If our ammunition were mathematically perfect, every round would exit the muzzle at the exact same point in that vibration wave.
  • But because of tiny, shot-to-shot manufacturing differences inside the box (described above in the “What Causes Group Size Variation section) each bullet exits the crown at a slightly different micro-second.
  • This causes them to leave the muzzle at slightly different launch angles. Two rounds with the same muzzle velocity but other differences do not always hve the same launch angle due to those other differences. Launch angle differences open up like a widening "V" over distance.
  • At 50 yards, this angular error is compressed and hard to see; at 100 yards, it expands into a major reason why your group opens up.
  • Lighter barrel launch angle variation is magnified because are more sensitive than heavier barrels.
2. MV variation increases vertical dispersion more and more as distance increases. When all else is the same, a 10 fps difference between rounds gives under 0.100" of vertical spread at 50; at 100 it's about 0.25". It's not linear.

3. Ammo flaws show more as distance increases: Small physical defects like a microscopic dent on a bullet's heel or center of gravity flaws don't cause much trouble at 50 yards. But as the bullet travels further and slows down, those tiny physical imbalances increasingly interfere with predictable trajectory.

4. Wind drift gets worse with distance: As distance grows, the effect of wind grows. A barely perceptible 1 mph crosswind can move a .22LR bullet about 0.10" at 50 yards but at 100 it’s three times as much.
 
What does this mean?

Judging a lot of ammo at 100 yards with 100 shots is like judging a baseball player by looking at his batting average only over a couple of games rather than over months or a season. You can’t be sure if you’re getting the right picture. Going hitless over two days or getting six hits are each only a snapshot of the a more complete picture. At 100 yards, one hundred shots isn’t sufficient to reliably capture a trustworthy picture of the ammo’s performance.

Using Extreme Spread (ES) or group size tracks only the two worst shots on the paper. A single group can instantly bloat up from a spectacular 0.80 inches to a terrible 1.40 inches just because one round had a very different launch angle, had a notable velocity drop, experienced a different puff of wind, or there was a minor change in rifle/rest interaction. It’s an inevitable weakness with ES as a metric.

This means that, when a single shot can really “make or break” a group at 100, more data is needed to help mitigate the problem as far as collecting reliable data is concerned. What was good at 50 isn't at 100.
 
The 100-Yard Breakdown

When you go out to 100 yards, the number of rounds you need to find the truth changes completely:
  • 10 Ten-Shot Groups (100 rounds / 2 boxes): This may be an internet standard for testing, but at 100 with .22LR it’s a statistical problem, a trap. At this distance 100 rounds is too small a sample size to average out the stacking variables. Different boxes from the exact same lot will give you wildly conflicting data, leading to false conclusions.
It is necessary to increase the total number of shots (and 10-shot groups). At 100 yards with .22LR much more reliable results require 30 ten-shot groups (300 rounds / 6 boxes). This is a much stronger baseline.

  • Increasing the sample size to 300 rounds provides enough data to more reliably overcome random anomalies.
  • It forces the data to account for a full range of manufacturing variations in the ammo, minor rifle set up inconsistencies, and changing range conditions, giving a reliable picture that doesn’t lie.
Many shooters might refuse to believe that two boxes of the exact same lot of match ammo can print completely different results in the exact same rifle. But it can. See the next section.
 
Part 3: Examples to Illustrate

In the third and final part, I am going to show data from a box-by-box experiment across six different rifles that illustrates exactly how deceptive a 100-round test can be. (And if it’s deceptive with 10-shot groups at 100, think about what it means about 5-shot groups at 50.)

Below is a chart showing results obtained at 100 yards with the same lot of .22LR Lapua Center X. It was shot over a number of sessions with six different rifles. Conditions were calm each but not necessarily equally calm each session. Each ten-shot group average was produced by two boxes of the ammo in question.

The first chart shows results for each rifle with ten 10-shot groups. The averages in blue show results indicating the rifle performed very well, that it “liked” the ammo. Averages in red show poor performance, that those rifles “hated” the ammo.

The 1907 and 2013 rifles “liked” the ammo; the 54.30 and 1913 “hated” the same ammo.

 
Here's a good summary video from 9-Hole reviews stating why 3 round groups suck, 5 round still aren't great, so they shoot 9 shot groups for maximum statistical accuracy

 
The second chart below shows the same rifles with the same lot of ammo but with different ten 10-shot groups. This time it’s the same lot, different boxes, and perhaps slightly different conditions (even if it’s the same day).

Here the 1907 and 2013 “hate” the same ammo. Now the 1913 and 1411 “like” the same ammo and the 54.30 seems to “like” it much more. It's a flip-flop. Who could have known?
 
Why the conflicting results?

The problem is simply that ten 10-shot groups are too small a sample size to be reliable for .22LR at 100 yards. When the shot count increases to at least 300, the picture becomes more reliable, more stable.

When each rifle shoots at least 300 rounds, the averages “settle” down to more reliable figures. It’s noteworthy, too, that the rifles shot the same ammo similarly.

This pattern can be repeated with other lots of .22LR match ammo tested. You may refer to more details in this thread https://www.rimfirecentral.com/thre...ads/a-comparison-of-multiple-lots-across-multiple-barrels-20000-rounds.1351053/

Ten 10-shot groups are too small a sample for reliability at 100 yards. A sample of 300 rounds in 10-shot groups is much more trustworthy.


 
What to takeaway from this thread?

Test with ten-shot groups. They are more useful and reliable than five-shot groups. The further the distance, the more shots are needed to get a reliable picture of the ammo.

 
5 shot groups are OK as long as you measure each one's size relative to the POA, not the groups center. You're just overlaying the same groups over each other in a way that makes it easier to see its trend.

It's also easier to mentally reset every 5-10 rounds then running a 25 rd banana mag. Like shooting a film camera, restrictions inherently make you slow down, focus, and act more intentionally.
 
Part 1 : Does your Rifle Really "Like" or "Dislike" the Ammo You’re Testing?

Every rimfire shooter soon understands that a group or two won’t tell him much of anything useful. That’s why he shoots a number of groups when he tests ammo.

The following, or some variation of it, happens every day at the range.:

"I did some serious lot testing today with two of my best rifles. I shot two full boxes of the same match ammo lot through each barrel in twenty 5-shot groups. After averaging the group sizes, Rifle A averaged a nice 0.350 inches, while Rifle B averaged a poorer 0..600 inches. Rifle A clearly loves this lot, but Rifle B does not."

He can honestly say “I shot the same ammo with two good rifles. One did well, the other didn’t.”

It sounds like a bulletproof, reliable experiment. The shooter didn't just fire one lucky group and call it a day; he put in the bench time, carefully shot 100 rounds in each rifle, measured every target, and calculated a mathematical average.

But despite all that hard work and experience, this shooter still did not get reliable information. It's a problem many shooters may experience.

On average, Rifle B doesn't "hate" the ammunition lot, and Rifle A doesn't have a magical preference for it. The problem usually isn't the shoooter or a lack of testing discipline -- the problem is the hidden mathematical trap built into the 5-shot group itself.
Carnage Carney says, "Who Cares"!...........:ROFLMAO::LOL::ROFLMAO:
 
If you shoot a full case of 5000 rounds, you'll have reliable data on whether the ammo is suitable for your purposes, but then you won't have any ammo left to use.

You can tell with two 5 round groups at 50 yards if an ammo is worth exploring further, a third group to confirm. I don't know the utility of focusing so much on ES, where the two most divergent rounds (for whatever causation) end up. Do the majority of the rounds land where you desire them, yes or no? How often do divergent rounds show up and how divergent are they? For your purposes, will a divergent round or two of this magnitude cause you to lose position in a match through no fault of your own? If yes, ammo unsuitable. If no, ammo suitable. This isn't so darn complicated, but I suppose some folks need something to do with idle hands.
 
Alternatively measure the distance to center for all ten rounds and use all 10 rounds instead of just two, one of which is most likely the least representative sample. Using group sizes skews the sample to towards the least representative rounds.
 
Mean radius is predictive. ES is not.

edit: And here's a spreadsheet that shows you the 90% confidence/error window for a given number of shots. The way to interpret this info is like so. You take your group size of a certain number of shots, multiply that by the value in the 0.05 column to get the lower end of the window, and multiply it by the value in the 0.95 column to get the upper end of the window. And that shows you how much uncertainty there is in that measurement. Say you're looking at a 5-shot group that measures 1 MOA, then the expected performance could lie anywhere between 0.603091 MOA and 1.435802 MOA. That's a pretty wide window. And that's why 5-shot groups don't really tell you much. 10 shots reduces that window to 0.73247-1.284208 MOA, which is still pretty large. I like to lot test with 100 rounds, which still gives a window of 0.91824-1.083378. The 300 shots Glenn mentions brings you down under +/- 5% with a window of 0.952524-1.047741. Takes about 7000 rounds to get that down under +/- 1%! Shooting is expensive, hehe.
 
Last edited:
Back
Top Bottom