Federal law does not ask whether you called it a test. Under the Uniform Guidelines on Employee Selection Procedures, a selection procedure is any measure used as a basis for an employment decision, and the definition reaches informal interviews and unscored application forms.
Your accounting candidate assessment therefore already exists, written down or not. The work is to turn it into a scored instrument you can compare across people, and to know which federal rules attach the moment it runs.
What an Accounting Candidate Assessment Has to Produce
Two things, or it has produced nothing: a score you can compare across candidates, and a decision rule you agreed to before you saw anyone's work.
The exercise itself is the easy half, and the seven-step hiring sequence already sets the task and the examples that go with it (how to hire a staff accountant). What decides whether any of it means anything is the rubric, a written scoring key that says what earns a mark and what it is worth, finished before the first file comes back.
Write the Rubric From the Job, Before You See Any Output
A review point is the unit. It is one thing a reviewer would send back: a misposted accrual, a reconciliation that ties only because a difference was plugged, a workpaper that cannot be retraced.
The federal guidelines describe the same discipline in their own vocabulary. A selection procedure can be supported by a content validity strategy, the argument that the exercise samples the job, to the extent that it is a representative sample of the content of the job, and the job analysis behind it should include an analysis of the important work behaviors required for successful performance and their relative importance (eCFR, section 1607.14C, Technical standards for content validity studies).
Written afterwards, a rubric is a description of the output you liked best.
Weight the Points by What the Error Would Have Cost
Counting mistakes treats a stray column width and a missed accrual as the same event. Sort review points into tiers instead, by what each error would have cost if nobody had caught it: cosmetic marks, rework inside the firm, and errors that reach the client. Give each tier a point value before you grade anything, or the tiers stay three piles instead of resolving into one comparable number. A candidate with several cosmetic marks and no client-facing errors is a different hire from one with a clean file and a single silent plug.
Set the Pass Condition and the Failure Action in Advance
Write down the score that passes and what happens to a candidate who does not reach it, before any output exists. The same point is already made at the agency level: a test with no failure condition is a demonstration with extra steps (how to judge a staffing agency).
At candidate level the failure action is a real decision with three honest options: reject, retest on a different task, or hire at a lower rung with the review that rung requires. Pick which one applies before you meet somebody you like. Then give every candidate for that seat the same task and the same rubric, because a comparison between two people who did different exercises is two opinions with numbers attached.
Every Selection Step Is Covered, Including the Casual Chat
The scope of the federal rules is wider than the word test suggests.
A selection procedure is any measure, combination of measures, or procedure used as a basis for any employment decision, and the term is written to reach the full range of assessment techniques, from traditional paper and pencil tests and performance tests through informal or casual interviews and unscored application forms (eCFR, section 1607.16Q, Definitions). The partner's chat at the end of the day sits in the same frame as the graded reconciliation. Scoring something does not pull it in. Using it to decide does.
Coverage is not unlimited, and it is worth settling before you build anything. The guidelines reach a user only to the extent it may be covered by federal equal employment opportunity law (eCFR, section 1607.16W, Definitions). Title VII defines an employer as a person engaged in an industry affecting commerce who has fifteen or more employees for each working day in each of twenty or more calendar weeks in the current or preceding calendar year (42 U.S. Code 2000e(b)). The disability rules run the same test, defining an employer as one with 15 or more employees for each working day in each of 20 or more calendar weeks in the current or preceding calendar year (eCFR, section 1630.2, Definitions).
Your own state's law sets its threshold separately. Below the federal line the rules that follow are not your duty, though your state's may be, and the discipline they describe is still the only way to compare two candidates on the same evidence. That is why a firm builds it before it has to.
Adverse Impact and the Four-Fifths Rule
Adverse impact is a substantially different rate of selection which works to the disadvantage of members of a race, sex, or ethnic group (eCFR, section 1607.16B, Definitions).
Use is where the duty bites. The use of any selection procedure which has an adverse impact on the hiring, promotion, or other employment or membership opportunities of members of any race, sex, or ethnic group will be considered discriminatory and inconsistent with the guidelines, unless the procedure has been validated in accordance with them, or the provisions of section 6 are satisfied (eCFR, section 1607.3A, Discrimination defined).
The rule of thumb has a precise wording. A selection rate for any race, sex, or ethnic group which is less than four-fifths, or eighty percent, of the rate for the group with the highest rate will generally be regarded by the federal enforcement agencies as evidence of adverse impact, while a greater than four-fifths rate will generally not be (eCFR, section 1607.4D, Information on impact).
It moves in both directions, and the same subsection says so. Smaller differences in selection rate may nevertheless constitute adverse impact where they are significant in both statistical and practical terms, or where a user's actions have discouraged applicants disproportionately on grounds of race, sex, or ethnic group. Greater differences may not constitute adverse impact where the differences are based on small numbers and are not statistically significant, or where special recruiting or other programs cause the pool of minority or female candidates to be atypical of the normal pool of applicants from that group (eCFR, section 1607.4D, Information on impact). Both of those alternatives, the discouraged applicant and the atypical pool, are what a small firm is most likely to meet, because how a job ad reads and where a referral push is aimed both shape who ends up applying.
Small numbers get their own treatment. Where the evidence indicates adverse impact but is based upon numbers which are too small to be reliable, evidence about the impact of the procedure over a longer period of time, or about its impact when used in the same manner in similar circumstances elsewhere, or both, may be considered (eCFR, section 1607.4D, Information on impact). A firm hiring a couple of people a year lives in that small numbers case, which is exactly why the record over several years is the thing worth keeping.
What You Are Expected to Keep
Each user should maintain and have available for inspection records or other information which will disclose the impact which its tests and other selection procedures have upon employment opportunities of persons by identifiable race, sex, or ethnic group (eCFR, section 1607.4A, Records concerning impact).
Small employers get a shorter version of that duty. To minimize recordkeeping burdens on employers who employ one hundred (100) or fewer employees, such users may satisfy the documentation requirements by keeping records that show, for each year, the number of persons hired, promoted and terminated for each job by sex and where appropriate by race and national origin, the number of applicants for hire and promotion on the same basis, and the selection procedures utilized (eCFR, section 1607.15A(1), Simplified recordkeeping for users with less than 100 employees).
That is a spreadsheet, not a compliance project, with one condition attached. If the user has reason to believe that a selection procedure has an adverse impact, the user should maintain any available evidence of validity for that procedure (eCFR, section 1607.15A(1), Simplified recordkeeping for users with less than 100 employees). The last column is the one that defeats a firm assessing every candidate on a different task, because there is nothing stable to name in it.
Keeping the sheet is also self-protection. Where the user has not maintained data on adverse impact as the documentation section of the applicable guidelines requires, the federal enforcement agencies may draw an inference of adverse impact of the selection process from that failure itself, if the user underutilizes a group in the job category compared with that group's representation in the relevant labor market, or, for jobs filled from within, in the applicable work force (eCFR, section 1607.4D, Information on impact).
What the ADA Changes About Testing
Four rules land directly on the assessment: two about what you may ask and when, one about accommodating the application process itself, and one about what the test is allowed to measure.
A medical examination cannot come first. Covered entity is the rules' own term for an employer, employment agency, labor organization or joint labor management committee (eCFR, section 1630.2(b), Definitions), so for a firm reading this it means the employer. Except as permitted by section 1630.14, it is unlawful for a covered entity to conduct a medical examination of an applicant or to make inquiries as to whether an applicant is an individual with a disability or as to the nature or severity of such disability (eCFR, section 1630.13(a), Prohibited medical examinations and inquiries).
The permitted version is timed and even-handed, and it sits beside a broad permission to ask about the work. An employer may require a medical examination after making an offer of employment and before the applicant begins their employment duties, and may condition the offer on the results, if all entering employees in the same job category are subjected to such an examination regardless of disability. It may also make pre-employment inquiries into the ability of an applicant to perform job-related functions, and ask an applicant to describe or to demonstrate how, with or without reasonable accommodation, the applicant will be able to perform them (eCFR, section 1630.14(b) and 1630.14(a), Medical examinations and inquiries specifically permitted).
Conditioning the offer on the results is not an unqualified right to withdraw it. The examination itself does not have to be job-related and consistent with business necessity, but if certain criteria are used to screen out an employee or employees with disabilities as a result of that examination or inquiry, the exclusionary criteria must be job-related and consistent with business necessity, and performance of the essential job functions cannot be accomplished with reasonable accommodation (eCFR, section 1630.14(b)(3), Medical examinations and inquiries specifically permitted).
Accommodation reaches the application process itself, since reasonable accommodation includes modifications or adjustments to a job application process that enable a qualified applicant with a disability to be considered for the position (eCFR, section 1630.2(o)(1)(i), Definitions), and not making it for an otherwise qualified applicant's known limitations is unlawful unless the employer can demonstrate undue hardship on the operation of its business (eCFR, section 1630.9(a), Not making reasonable accommodation).
The test also has to measure what it claims to. It is unlawful to fail to select and administer employment tests in the most effective manner to ensure that, when a test is given to an applicant whose disability impairs sensory, manual or speaking skills, the results accurately reflect what the test purports to measure rather than the impaired skills, except where those skills are the factors the test purports to measure (eCFR, section 1630.11, Administration of tests).
In practice that is one line in the invitation saying how to ask for an adjustment, plus one decision made once about which parts of the exercise are the measurement and which are incidental. A clock measures speed. If the seat does not need speed, the clock is measuring the wrong thing for everybody.
What to Assess at Each Rung
Three instruments do the work at any rung: a sample of real work, a structured interview, and a reference call. What changes from rung to rung is what each one has to reveal. The rung sets the test, and your review capacity sets the rung.
The mapping runs from the seat you are filling to the thing the exercise has to surface, and then to the person who has to be capable of marking it.
| Rung | What the exercise has to reveal | Who has to be able to grade it |
|---|---|---|
| Bookkeeper or accounts payable | Whether the item that does not fit gets flagged or quietly forced | Whoever reviews that work in your firm today |
| Staff accountant | Whether the work can be reviewed at all, or has to be re-performed | The person who reviews staff workpapers |
| Preparer | What happens to the missing piece, and whether the question gets raised | Somebody who already reviews returns |
| Reviewer | Judgment, not detection alone | A partner, or the reviewer above that seat |
What the first three rungs need from a sample, and the exercises that draw it out, are already written up in how to hire a staff accountant. The rung that sequence does not reach is the reviewer, and it is the one where the grading changes shape. Hand over somebody else's finished file, score what gets caught, then score separately what goes back to the preparer, what gets fixed in place, and what gets escalated.
Testing above the rung is the expensive mistake. A senior exercise handed to every applicant selects for people you will then park on reconciliations it never covered. The real ceiling is your own bench, because the highest rung you can hire is the highest rung somebody in the firm can grade.
The Layers a Work Sample Cannot Reach
Two layers sit outside the sample, and both take the same scoring discipline or they take none.
A structured interview measures job-related competencies by systematically inquiring about a candidate's behavior in past experiences or their proposed behavior in hypothetical situations, and it runs on the rubric's discipline. All candidates are asked the same predetermined questions in the same order, and all responses are evaluated using the same rating scale and standards for acceptable answers (US Office of Personnel Management, Structured Interviews).
Reference calls take the same shape or they are worth nothing. Ask about the work behaviors your rubric names instead of asking for a general impression, put the same questions to every referee, and record the answers against the same scale. A referee who cannot speak to a behavior at all has told you something.
One line of law attaches the moment you buy a report instead of making the call yourself. Using a company in the business of compiling background information brings Fair Credit Reporting Act duties before you order it and again before you act on it, set out step by step in the hiring sequence (how to hire a staff accountant).
Assessing a Candidate You Did Not Source
An agency or an offshore provider hands you a person and a summary of how good that person is. The summary is the vendor's instrument, scored against the vendor's rubric.
Run your own block of real work at the person level, with the individual's name attached to the file, graded by your reviewer against your rubric, and keep the score beside the scores of the candidates you sourced yourself. That answers whether this person can hold this seat, which is a different question from whether the provider is any good.
Where the block comes from decides whether you can build it at all. A test block made of live client return information is a disclosure of that information whichever side of the border the candidate sits on. The route that carries no taxpayer consent is a narrow one. A tax return preparer may disclose return information to another tax return preparer located in the United States, including any territory or possession of the United States, for the purpose of preparing or assisting in preparing a return, or obtaining or providing auxiliary services in connection with that preparation, so long as those services are not substantive determinations, meaning an analysis, interpretation or application of the law, or advice affecting the tax liability reported (eCFR, section 301.7216-2(d)(1), Preparer-to-preparer disclosures).
A hiring exercise is not return preparation, which is why the block gets built from files you are already cleared to disclose. Where the candidate sits outside the United States, signed taxpayer consents are the step in front of that, and the sequence is already set out (how to judge a staffing agency).
When Not to Run One
Do not run an assessment you cannot grade this week, because an exercise that sits in an inbox until the candidate accepts another offer costs you the candidate and teaches you nothing. Do not run a long one where the cost of being wrong is low and the fix is quick, since the instrument exists to price the risk in the decision and your own records can price a wrong hire directly (the cost of a bad hire). And do not run one at all until you know which seat you are filling, which is the decision upstream of this one (accounting firm hiring strategy).
Write the Rubric Before the Next Candidate
Take the exercise you already send and finish the missing half. Name the review points, weight them by what each error would have cost, write the score that passes, and write what happens to somebody who does not reach it. Then keep the sheet, because it is both the comparison and the answer to how you decided.
If the seat cannot wait for a search to finish, the same discipline works on bought capacity. Accountably places trained offshore accountants and tax preparers inside US CPA and EA firms, ramped on the firm's own software and procedures in about 3 to 4 weeks, with 30+ placements across 20+ firms since 2022.
The Free 40-Hour Proof Pilot puts a placement through the rubric you just wrote. A fixed block of your own representative work, prepared on your procedures and put through the full review chain, so your reviewer grades real output against your own standard before anything is riding on it. If a placement is not the right fit in the first 30 days, we replace them free. Don't trust us. Test us.
