Key Takeaways
- 63% of employers say candidates increasingly base salary requests on inaccurate or unverified data. The numbers you bring to a negotiation determine whether you leave money on the table or ask for a number that gets the offer pulled (Payscale).
- Self-reported salary platforms like Glassdoor skew 5% to 15% above true market medians. The bias is predictable: above-average earners report more, and active job seekers are either underpaid and leaving or overpaid and gaming their next offer.
- 73% of salary negotiations fail not because candidates ask for too much, but because they anchor their requests to fundamentally flawed data. The source of your number matters more than the number itself.
- The three-source rule eliminates most data quality problems. Cross-reference a government dataset for the baseline, a role-specific platform for current ranges, and a job posting aggregator for employer-disclosed pay before you name a single number.

Why Is Most Online Salary Data Less Accurate Than You Assume?
Most job seekers approach salary research the same way they approach any other online search. They type a job title and location into the first platform that appears, glance at the average, and treat that number as fact. The process takes ninety seconds. The consequences of getting it wrong compound over an entire career, because each salary anchors the next one and a 5% shortfall in your first negotiation is not a one-time loss. It is a compounding one.
The structural problem with online salary data is not that it is intentionally misleading. It is that the incentives of the platforms that collect it are fundamentally misaligned with the accuracy needs of the people who use it. Self-reported salary sites depend on user submissions to build their databases. The users who submit are not a random sample of the workforce. They are disproportionately people who earn above average and want to share, or people who earn below average and want to benchmark before leaving. Both groups introduce sampling bias, and both push the aggregate numbers away from the true market median. Research from Stanford labor economists comparing self-reported platforms against verified employer data found that Glassdoor salaries skew 5% to 15% above actual market medians (Payscale, "Compensation Best Practices Report", 2025). The direction of the error is consistent. The magnitude varies by role. The implication is the same in every case: the number you are looking at is probably higher than what most employers actually pay.
Compounding the sampling problem is the staleness problem. Salary data on free platforms can sit unrefreshed for years. A software engineer salary from 2022 reflects a hiring market that no longer exists. A marketing manager salary from 2023 predates the AI-driven restructuring that has changed what companies pay for those skills. Relying on data that has not been updated since the last presidential administration is not research. It is gambling with a number that shapes your financial trajectory, and the odds are not in your favor.
73% of salary negotiations fail because candidates anchor their requests to fundamentally flawed market data, not because they ask for too much. The data source determines the outcome. A candidate who walks into a negotiation with a number from a verified, employer-reported dataset makes a different impression than one who cites a Glassdoor average they looked up ten minutes before the call. Hiring managers can tell the difference. Compensation teams can certainly tell the difference. (Payscale, "Compensation Best Practices Report", 2025)
The Salary Data Landscape: Where the Numbers Actually Come From
Before you can choose the right data source, you need to understand what kind of data each source actually contains. Salary platforms do not all measure the same thing. They do not all collect data the same way. And they do not all serve the same purpose in a negotiation. Grouping them by data type rather than by brand name clarifies which ones belong in your research stack.
Government datasets sit at the top of the reliability hierarchy. The Bureau of Labor Statistics Occupational Employment and Wage Statistics program surveys approximately 1.1 million establishments across 830 occupations and reports actual employer-submitted wage data with geographic breakdowns. The data is statistically rigorous, covers nearly every industry, and carries the weight of a federal mandatory survey. It also lags the market by twelve to eighteen months and does not include bonuses, equity, or benefits, which together add an average of 29.5% to total compensation. Government data gives you the floor. It does not give you the ceiling, and it does not give you the specific number a tech startup in Austin will offer a senior product manager next week (Salary.com).
Crowdsourced platforms like Levels.fyi and Glassdoor fill in the timeliness gap that government data cannot close. Levels.fyi, which has become the standard reference for technology compensation, verifies submissions against offer letters and W-2 forms rather than accepting self-reports at face value. It breaks total compensation into base salary, equity, signing bonus, and performance bonus, which matters because equity can represent 40% or more of total comp at public tech companies. The tradeoff is coverage. Over 70% of Levels.fyi submissions are software engineering roles. The platform is exceptional for its niche and nearly useless for anyone outside it.
Job posting aggregators represent the newest category and, in some ways, the most directly useful one for job seekers. Platforms like Comprehensive.io and HireJack extract salary ranges directly from employer job postings rather than relying on employee self-reports. Because pay transparency laws now cover roughly 60 million U.S. workers across sixteen states, an increasing share of postings include employer-disclosed salary ranges. The data is as current as the job market itself because it comes from live postings. The limitation is coverage. If a role is not being actively recruited with a disclosed range, it will not appear in the dataset. But when it does appear, the number comes from the employer's own budget rather than a stranger's self-report.
The numbers in the chart reflect what happens when salary research goes wrong, and how often it does. More than half of employees never negotiate at all, leaving whatever is on the table untouched. Nearly two-thirds of employers report candidates arriving with unverified data. Almost half of all self-reported salary figures are off by more than 15%. And nearly three-quarters of failed negotiations trace back to weak data, not weak candidates. The pattern is not subtle. The people who lose negotiations are not less qualified. They are less informed, and they are less informed because they trusted the easiest source instead of the right one.
Which Sources Should You Trust for Your Specific Situation?
No single salary platform is correct for every role, industry, and negotiation context. The right source depends on what kind of role you are researching, what stage of the hiring process you are in, and what kind of company you are talking to. The matrix below maps the major data source categories against the scenarios where each one earns its place in your research.
| Data Source | Best For | Watch Out For |
|---|---|---|
| BLS / Government Data | Establishing the defensible floor for any role; geographic comparisons between metro areas; industries where government data is comprehensive such as healthcare, education, and manufacturing | 12 to 18 month lag; no equity, bonus, or benefits data; occupational categories are broad and may not map cleanly to specific job titles |
| Levels.fyi (Verified) | Technology roles at companies with established compensation bands; roles where equity is a significant portion of total pay; comparing offers between tech employers | Over 70% of submissions are engineering roles; skews toward higher offers and big tech; limited value for non-tech industries |
| Glassdoor / Payscale (Self-Reported) | Broad industry coverage when government data is too stale; company-specific salary research; getting a directional sense before deeper research | Self-report bias inflates numbers 5% to 15% above true median; data can be years old; treat as directional, not definitive |
| Job Posting Aggregators | Current market rates for actively recruited roles; verifying whether an offer aligns with what competitors are advertising; roles in states with pay transparency laws | Only captures roles being actively recruited; gaps where disclosure is not legally required; U.S.-centric coverage |
| Industry Salary Guides (Robert Half, Randstad) | Roles in finance, accounting, legal, and administrative functions where these guides have deep historical data; getting percentage change year-over-year for a function | May reflect placement data skewed toward larger metro areas; typically annual updates; limited granularity for niche roles |
The table is not a ranking. It is a decision tool. A software engineer negotiating an offer from a Series C startup needs Levels.fyi for the equity data and a job posting aggregator for what competitors are advertising. A nurse researching pay in a specific hospital system needs BLS data for the regional baseline and Glassdoor for facility-specific ranges. Different roles, different stacks, different numbers. The only universal rule is that one source is never enough.

How to Triangulate: The Three-Source Rule That Eliminates Bad Data
The single most reliable Salary Research habit is triangulation: checking three independent data sources and only trusting numbers that converge within a reasonable band. If BLS says the median for your role in your city is $72,000, a job posting aggregator shows current postings ranging from $68,000 to $82,000, and Levels.fyi shows verified offers clustering around $75,000, you have a narrow, defensible range. If Glassdoor says $95,000 and everything else says $72,000, you have an outlier to discard, not a bargaining chip.
The sequence matters as much as the count. Start with government data for the baseline. It is the slowest to move but the hardest to dismiss because it is statistically representative in a way that no self-reported dataset can match. Then check a role-specific source: Levels.fyi for tech, a professional association survey for law or medicine, an industry salary guide for accounting or administrative roles. Finish with a current-market check using a job posting aggregator or by searching live postings in your target geography and filtering for roles that include disclosed salary ranges. The whole process takes twenty minutes. The difference between a triangulated number and a single-source guess is the difference between negotiating from strength and hoping you did not just ask for something that ends the conversation.
When you have your range, attach a confidence level to it. A number confirmed by government data, verified platform submissions, and current employer-posted ranges is high confidence. A number from a single self-reported platform with wide variance and stale entries is low confidence. You do not need to share your confidence rating with the employer. You need to know it yourself so that you anchor your ask to the high-confidence data and hold the low-confidence data loosely enough to adjust when the employer's offer reveals information you could not have gathered in advance.
From Research to Negotiation: Using Data Without Sounding Like a Spreadsheet
Gathering accurate Salary Data is the prerequisite. Delivering it in a negotiation without sounding robotic or adversarial is the skill that converts research into dollars. Most candidates make one of two mistakes. They either cite their data like they are presenting evidence in court, which puts the employer on the defensive, or they fail to cite any data at all, which leaves their ask floating in the air with nothing underneath it. The middle path works better than either extreme.
Open with the range rather than a single number. A single number invites a yes or no. A range opens a conversation. The phrasing that works is some version of: "Based on my research into current market rates for this role in this geography, I was expecting something in the range of X to Y. How does that compare with what you have budgeted?" The question at the end turns a demand into an invitation. It also gives the employer room to say that your range is within their budget, slightly above it but negotiable, or far above it, in which case you learned something valuable about the role before accepting it.
Reference your sources in passing, not in detail. You do not need to name Levels.fyi or BLS by name in a negotiation any more than you need to cite your sources in a conversation about the weather. What matters is that you sound like someone who did the work, not someone who is guessing. A phrase like "market data for comparable roles in this region" communicates research confidence without sounding like you are reading from a deposition. If the employer pushes back on your range, that is when you get specific. "I am looking at verified offer data for this level at comparable companies, plus current postings that disclose ranges between X and Y." Specificity is your backup. Confidence is your opener.
The candidates who win negotiations are not the ones with the most aggressive asks. They are the ones whose asks are built on numbers that both sides can treat as real. An employer can argue with your preferences. It is much harder for them to argue with verified market data presented calmly, without ego, as a shared starting point for a conversation that both parties want to end with a yes. The research does the arguing for you. Your job is to deliver it like it is the most natural thing in the world, because after twenty minutes of triangulation, it is.
Accurate salary data does not guarantee a higher offer. No data can force an employer to pay more than their budget allows, and no source can predict how urgently they need to fill the role. What accurate data does is eliminate the single largest unforced error in salary negotiation: asking for the wrong number because you trusted the wrong source. Fixing that error costs nothing except the twenty minutes it takes to check three sources instead of one. Over a career, the return on those twenty minutes is measured in hundreds of thousands of dollars. There are not many investments with that ratio. There are even fewer that almost nobody makes.