A polo graded correctly from a 4 to a 16 should feel like the same shirt on every body it fits. Golf apparel rarely pulls that off. The swing raises the stakes on a fit miss: a shoulder seam that sits half an inch wrong doesn't just look off, it restricts the motion the whole garment was built to support.
Most sizing problems don't start on the cutting table. They start in the grade rule, the small set of measurements that decides how much a garment grows or shrinks between sizes, and where on the body that growth happens.
Grading isn't scaling. It's a series of decisions.
A grade rule is not "add an inch and repeat." Different parts of the body grow at different rates between sizes, and a garment has to follow that curve or it stops fitting like the sample everyone approved.
The chest grades faster than the waist. The shoulder barely grades at all. A skort's rise grades differently than its hem. Golf sharpens this problem past what most sportswear deals with: a shoulder seam graded like a T-shirt's will bind on the backswing in every size above the sample.
Most brands approve one size, usually a middle size, and assume the grade rule carries the fit through the rest of the range. It often doesn't. Nobody tries on the smallest or the largest size before the run ships.
The sample size lies about the extremes.
A size 8 sample that moves cleanly through a full shoulder turn tells a brand almost nothing about how a size 18 or a size 2 will move. The grade rule is doing all the work at the far ends of the range, and it is rarely the part that gets tested.
This is where inclusive sizing actually breaks. Not in the decision to offer more sizes, but in the unglamorous work of proving the grade holds up once a brand does. Extending a range without testing the extremes just moves the fit problem further out and hands it to whoever returns the garment.
A second fitting, on a piece pulled from the real production run rather than the approved sample, is the only way to know for certain. Few programs budget the time for it. Fewer still budget the willingness to fix a grade rule after tooling is already cut.
Fabric behavior complicates the math further.
A four-way stretch knit and a woven bottom don't grade the same way, even inside the same collection. Stretch fabric can hide a grading error standing still and expose it in motion, which is exactly the moment a golfer needs the garment to disappear.
Construction choices interact with grading in ways that are easy to miss on paper. A princess seam and a dart do different things to how a chest measurement translates into fabric. Change the construction between sample rounds and the old grade rule may no longer apply, even if nobody touched a single number.
The fix happens before cutting, not after the complaints.
By the time return data shows a size running small, the grade rule has already shipped through however many SKUs were in that collection. Fixing it after the fact means re-grading, re-cutting, and re-approving a range that customers have already formed an opinion about.
The cheaper version of this work happens at the tech pack stage: grade rules specified by pivot point, not just by size, and checked on more than the sample size before a run gets committed. It is slower. It is also the only point in the process where a fit problem costs a fitting instead of a return.
A brand can promise performance and design in the same collection. It cannot promise both in a grade rule nobody checked past the sample.