There are multiple reasons. First, the tabulator typically needs to determine the "style" (layout) of the ballot from one of possibly hundreds of options (due to multiple government and taxation jurisdictions). Second, a typical ballot in most US jurisdictions has a dozen or more questions on it, and cumulatively across all ballot styles there are often over a hundred questions.
Third, though, and perhaps most significantly, "traditional" optical mark reading (OMR) systems using LED or laser sources and diodes were inflexible as to ballot layout and more problematically not very reliable across varying marks (remember the grade-school requirement for #2 pencils due to OMR scoring of exams), a particularly big issue since voters are often not experienced with OMR systems and so do not mark their ballot "correctly." To address this, almost all modern ballot tabulators use a CCD mechanism to take an image of the full ballot and then interpret it via machine vision (this is not a case of machine learning, the algorithms used are actually very simple). This yields much more reliable interpretation of ballots with fewer ballots rejected to hand-counting, but requires more complex software.
It's important to understand that most US election administrators avoid hand-counting in large part because of its inaccuracy. In many US jurisdictions hand-count ballots are counted by two individuals to improve reliability, but the error rate remains higher than machine tabulation. When it is 1AM after a day that started at 5AM and you are on the hundredth ballot you've hand-tabulated since you got off the precinct floor it becomes extremely difficult to tabulate with the virtually zero error rate that US voters expect. This is not a hypothetical scenario but one that's pretty typical of US election working conditions due to the slim budget and expectation of rapid posting of returns.
Third, though, and perhaps most significantly, "traditional" optical mark reading (OMR) systems using LED or laser sources and diodes were inflexible as to ballot layout and more problematically not very reliable across varying marks (remember the grade-school requirement for #2 pencils due to OMR scoring of exams), a particularly big issue since voters are often not experienced with OMR systems and so do not mark their ballot "correctly." To address this, almost all modern ballot tabulators use a CCD mechanism to take an image of the full ballot and then interpret it via machine vision (this is not a case of machine learning, the algorithms used are actually very simple). This yields much more reliable interpretation of ballots with fewer ballots rejected to hand-counting, but requires more complex software.
It's important to understand that most US election administrators avoid hand-counting in large part because of its inaccuracy. In many US jurisdictions hand-count ballots are counted by two individuals to improve reliability, but the error rate remains higher than machine tabulation. When it is 1AM after a day that started at 5AM and you are on the hundredth ballot you've hand-tabulated since you got off the precinct floor it becomes extremely difficult to tabulate with the virtually zero error rate that US voters expect. This is not a hypothetical scenario but one that's pretty typical of US election working conditions due to the slim budget and expectation of rapid posting of returns.