هذا مخرَج حقيقي غير مُحرَّر من ProSearch، وُلِّد من مصادر أكاديمية في دقائق. وكل [Ref N] يقابل مصدراً في مكتبتك.
[Strategic Research Advisory: Impact of Artificial Intelligence]
Abstract
Research on the impact of Artificial Intelligence (AI) has expanded rapidly across economics, finance, accounting, auditing, marketing, entrepreneurship, and law, yet the field still suffers from a recurring “impact-evidence gap”: expectations of transformative productivity and decision-quality improvements often outpace what is visible in conventional statistics and causal empirical designs. Early macro explanations emphasize implementation lags, complementary intangible capital, diffusion frictions, and mismeasurement as key reasons aggregate productivity can remain sluggish even while AI capabilities surge [Ref 1]. Recent firm-level synthesis work argues that “AI” is not a single measurable object—datasets variously capture invention vs. use, internal capability building vs. outsourcing, realized activity vs. investor perceptions—so empirical conclusions will differ unless measurement choices are explicitly matched to the research question [Ref 21].
Within business functions, evidence is strongest where researchers observe concrete adoption signals and outcomes. For auditing, firm AI hiring is associated with measurable improvements in audit quality (lower restatement likelihood), lower fees, and delayed labor displacement effects, supported by both large-scale data and partner interviews [Ref 7]. In financial reporting and earnings-related contexts, emerging work shows that how firms *talk* about AI (e.g., AI term frequency in annual reports) correlates with certain earnings management patterns, especially real activities manipulation, though effects vary by earnings management type and proxy choice [Ref 5]. Parallel text-as-data approaches applied to earnings calls show a sharp post-ChatGPT rise in AI mentions and increased uncertainty language, enabling near-real-time measurement of firm sentiment and perceived risk [Ref 20]. Meanwhile, LLM-based approaches to extracting “core earnings” from 10-Ks show strong dependence on prompt structure: “out-of-the-box” use fails on complex accounting reasoning, while sequential prompting performs markedly better and can be cost-effective at scale [Ref 18].
The most valuable gaps are (i) causal identification of AI’s real operational impacts versus disclosure/“AI washing,” (ii) harmonized measurement linking AI capability-building to specific process changes, (iii) multi-modal disclosure analytics (text+audio+video) as earnings calls evolve toward videocasts [Ref 10], and (iv) governance, ethics, and regulatory design for responsible AI deployment in market-facing functions [Ref 4]. Top recommendations are: build validated AI adoption indices combining workforce, disclosure text, and vendor/outsourcing signals; run quasi-experimental designs around adoption shocks; and develop multi-modal models that connect managerial communication cues to subsequent financial reporting quality, audit outcomes, and market reactions.
Keywords
Artificial intelligence impact; productivity paradox; AI measurement; firm-level adoption; earnings management; audit quality; large language models; text-as-data; earnings calls; AI governance; intangible capital; causal inference
1. Introduction
Artificial Intelligence (AI) is widely recognized as a general-purpose technology with the potential to reshape production, decision-making, labor demand, and market structure. In many domains, AI systems now match or surpass human performance on narrowly defined tasks, fueling optimism about productivity and economic welfare. Yet a central tension persists: measured aggregate productivity growth has been weak in many advanced economies even as AI capabilities and investment surge [Ref 1]. This mismatch matters scientifically (it challenges how we conceptualize technology diffusion and measurement) and practically (it affects policy, firm strategy, and workforce planning).
Research has established that technological revolutions often require complementary innovations—organizational redesign, new skills, process reengineering, data infrastructure, and governance—before productivity benefits appear broadly. This implies that the main “impact” of AI may not be immediate and may be poorly captured by standard national accounts. Brynjolfsson, Rock, and Syverson explicitly frame these complementary investments as intangible capital and argue implementation lags are likely a leading explanation for the modern productivity paradox [Ref 1].
At the micro level, the literature is increasingly shifting from “Does AI matter?” to “Which AI efforts matter, for whom, under what conditions, and how can we measure them credibly?” A recent synthesis emphasizes that AI measurement choices are not secondary technicalities: different datasets capture different underlying economic objects (invention vs. use; internal vs. outsourced; realized activity vs. perceptions), leading to different empirical conclusions. The purpose of this advisory is to critically assess your provided sources, locate them within the broader research landscape, identify high-value research gaps, and propose publishable, methodologically credible research directions—especially those that can survive peer review in high-impact economics, finance, accounting, information systems, and management outlets.
2. Critical Analysis of Reviewed Sources
2.1 Artificial Intelligence and the Modern Productivity Paradox: A Clash of Expectations and Statistics
This working paper articulates the core puzzle: AI performance progress and soaring market valuations coexist with slowed measured productivity and stagnant incomes for many Americans [Ref 1]. It proposes four explanations—false hopes, mismeasurement, redistribution, and implementation lags—and argues lags and complementary innovations are likely central. Its strength is conceptual clarity and linkage to general-purpose technology history; its limitation is that, based on the provided excerpt, it is more framework-building than providing definitive causal estimates. It leaves open the micro-to-macro linkage: which firm-level AI investments translate into measurable productivity and when.
2.2 Artificial Intelligence and Corporate Earnings Management
This study uses survey-style quantitative data from 145 SMEs in Lagos State, Nigeria and applies PLS-SEM with bootstrapping (5000 replicates) to test relationships between AI-related constructs and corporate earnings management [Ref 2]. It reports very high explanatory power (R² = 0.990) and statistically significant correlations for five hypothesized drivers, with “integration of AI with existing systems” showing the largest correlation. The methodological clarity (PLS-SEM with reliability/validity metrics reported) is a strength; however, the extremely high R² raises concerns about common-method bias, construct overlap, and overfitting risk in perceptual measures. The study is also context-specific (SME manufacturing in one region), which limits generalizability.
2.3 The Impact of Artificial Intelligence and Blockchain on the Accounting Profession
This paper provides a broad review of AI, machine learning, big data, and blockchain in accounting practice and education, highlighting changing skill demands and governance concerns [Ref 3]. It lists AI technology categories (e.g., NLP, ANN, computer vision) and many named applications, and it frames implications for educators and professionals. Its strength is integrative coverage and concrete mapping from technologies to accounting scenarios; its weakness is that it appears primarily descriptive/review-based rather than providing systematic evidence on outcomes like reporting quality, audit quality, or earnings management.
2.4 Leveraging Artificial Intelligence in Marketing for Social Good—An Ethical Perspective
This work systematically scrutinizes ethical challenges of deploying AI in marketing from a multi-stakeholder perspective and highlights tensions between ethical principles, questioning a purely deontological approach [Ref 4]. It contributes a governance lens and proposes how AI can be leveraged for societal/environmental well-being. While not finance/accounting-specific, it is highly relevant for AI impact research because disclosure, persuasion, targeting, and stakeholder effects are core mechanisms through which AI changes markets. The limitation is that it offers ethical analysis and suggestions rather than operational metrics or causal tests.
2.5 Research On the Impact of Artificial Intelligence on Corporate Earning Management
Using panel data from A-share non-financial listed firms (2005–2024), this study measures AI-related discourse frequency in annual reports and relates it to real activities manipulation (RM) and accrual-based earnings management (AM) via fixed-effects regressions [Ref 5]. Its key contribution is a scalable text proxy for “AI embeddedness” and a clear distinction between RM and AM, finding stronger association for RM than AM. The main limitation is construct validity: AI term frequency may capture disclosure strategy or hype rather than actual adoption, and may be endogenous to governance quality and investor relations incentives.
2.6 Artificial Intelligence and Entrepreneurship: Implications for Venture Creation in the Fourth Industrial Revolution
This article conceptualizes how AI augments/replaces tasks in venture processes (idea production, selling, scaling) and advances a research agenda emphasizing both benefits and liabilities, including risks of disintermediation for traditional small firms [Ref 6]. It is valuable as a theory-forward agenda-setter that points to distributional and organizational design implications. Its limitation for empirical publication is that it does not, in the provided text, specify testable operationalizations or datasets—so follow-on empirical work is needed.
2.7 Is artificial intelligence improving the audit process?
This study combines a large dataset of resumes (over 310,000) to identify audit firms’ employment of AI workers with interviews of 17 audit partners [Ref 7]. It finds AI investment is associated with lower restatement likelihood (5.0% reduction per one-standard-deviation change), lower fees (0.9% drop), and gradual displacement of accounting employees over 3–4 years. Its strengths are measurement innovation (AI workforce), mixed-method triangulation, and outcome-based evidence. A remaining gap is mechanism unpacking: which audit tasks, tools, and governance practices drive quality gains versus cost cutting.
2.8 Forecasting Earnings, Artificial Intelligence (AI) Versus Equity Analysts
This source summarizes a study using LLMs to estimate core earnings from 10-K filings, contrasting a minimal “lazy” prompt with a structured sequential prompting approach, and reporting comparative forecasting performance metrics and error statistics. Methodologically, it is detailed about prompting workflows (multiple calls) and evaluation logic; however, it is a secondary summary rather than the primary paper itself, so for academic publication you would need to cite and engage the underlying study directly. It also highlights a key research problem: LLMs can be fragile without guidance.
2.9 Twenty-five years (1998-2023) of Earnings Disclosure with Conference Calls: a Bibliometric Review focusing on Artificial Intelligence
This bibliometric review analyzes 14,437 articles (1998–2023) and identifies themes in earnings call research, emphasizing corporate governance and categories like capital markets and sustainability [Ref 9]. It explicitly proposes future AI-enabled research using advanced text, voice, and image analysis, particularly as calls evolve toward videocasts. Its strength is mapping the intellectual landscape and motivating multi-modal AI; its weakness is that bibliometrics does not establish causal relationships or validate specific AI measures.
2.10 Twenty-five years (1998-2023) of Earnings Disclosure with Conference Calls: a Bibliometric Review focusing on Artificial Intelligence
This appears to be the same work as Ref 9 with identical abstract and key points [Ref 10]. Treat it as a duplicate in your bibliography management. It reinforces the same gap: opportunity for AI in analyzing text, audio, and image signals in videocast-style disclosures.
2.11 The Global Artificial Intelligence Revolution Challenges Patent Eligibility Laws
This article discusses how AI progress challenges patent eligibility laws and illustrates AI applications across domains, including claims about productivity gains from deep learning-based quality inspection. It is useful for understanding legal/institutional constraints shaping AI diffusion and commercialization. The limitation is that it is not designed as an empirical economic impact study; many examples are illustrative, and you should be cautious about treating them as generalizable evidence.
2.12 Firm Maturity, AI Adoption, and Financial Performance: A Study of Publicly Listed Service Firms in Finland
This master’s thesis uses annual report keyword frequency to measure AI adoption and relates it to ROA/ROE, considering firm maturity and using OLC and TOE frameworks [Ref 12]. Its contribution is an explicit theoretical framing and a practical measurement approach. The limitation is common to many keyword-frequency adoption proxies: mentions may reflect strategy signaling rather than implementation, and performance measures may lag adoption. As a thesis, it may also face data and design constraints compared to journal standards.
2.13 AI-Driven Valuation Techniques in Cross-Border Mergers and Acquisitions: Enhancing Accuracy in Emerging Markets
This study compares a Random Forest Regressor approach to a conventional EV/EBITDA valuation method for cross-border M&A in emerging markets, arguing AI improves predictive accuracy and robustness under data inconsistencies and volatile conditions [Ref 13]. The strength is a clear “AI vs traditional” benchmarking framing with a specific model; the weakness is that valuation “accuracy” depends heavily on ground truth definition (deal value? post-merger outcomes?) and on data representativeness. It is promising but would benefit from clearer causal and decision-use validation.
2.14 The Impact of Artificial Intelligence on Accounting Information and Earnings Management: Bibliometric Analysis
This RePEc record provides the title, author, and abstract-level metadata indicating a bibliometric analysis on AI’s impact on accounting information and earnings management [Ref 14]. Because the full text is not included here, you should treat it as a lead for retrieval rather than a source of detailed findings. It signals growing scholarly consolidation in this niche, but you will need the paper to extract methods, datasets, and specific gaps.
2.15 Section 03. Smart Solutions in IT — “Intuitive Intelligence”
This short piece discusses “intuitive intelligence” and claims about job shares related to computing, alongside examples of AI-assisted design [Ref 15]. Its contribution is limited for academic impact research: it is speculative, lacks methodological grounding, and includes broad claims that would need verification. Use it, if at all, only as contextual narrative, not as evidentiary support.
2.16 Tackling Climate Change with Machine Learning
This work outlines how ML can help reduce emissions and support climate adaptation, identifying high-impact problems and research gaps [Ref 16]. While not directly about corporate finance/accounting, it is valuable as an exemplar of “AI for social good” research framing and gap identification. The limitation is topical distance: its main utility here is methodological inspiration for impact-oriented research programs.
2.17 How AI is Transforming Earnings Report Processing in 2025
This piece describes AI-driven extraction, reconciliation, and compliance checking in earnings report processing and includes quantitative performance claims (e.g., accuracy rates, time reductions) tied to a specific company’s tools [Ref 17]. It provides concrete process-level hypotheses and potential outcome measures; however, it is not presented as peer-reviewed research and contains marketing-style claims that require independent validation. Treat it as practitioner context and a source of testable propositions, not as definitive evidence.
2.18 AI-Powered Core Earnings Analysis: A New Frontier in Financial Reporting
This source summarizes findings from “Scaling Core Earnings Measurement with Large Language Models,” emphasizing that structured prompting improves LLM performance and reporting specific predictive power comparisons and cost-per-firm claims. It is useful for motivating research on prompt engineering, reliability, and scalable accounting measurement. As with Ref 8, it is a secondary source; for publication you should obtain and cite the underlying study directly.
2.19 Using Data Analytics and Artificial Intelligence for Public Disclosures
This piece provides guidance on using data analytics and AI to manage compliance risks in public financial and sustainability disclosures, including risk detection, activism/litigation response, fraud, and misinformation [Ref 19]. It is valuable for mapping practical governance and risk domains and for motivating measurable compliance outcomes. However, it is advisory in nature, so academic contribution requires translating recommendations into testable hypotheses and evaluable interventions.
2.20 AI Optimism and Uncertainty: What Can Earnings Calls Tell Us Post-ChatGPT?
This analysis applies a text-as-data approach to a very large corpus of earnings call transcripts (185,999 calls; 7,047 U.S. firms; 2008–2024Q1), documenting a more than fivefold rise in AI-mention sentences post-ChatGPT and measuring sentiment and risk/uncertainty language using a finance-specific lexicon [Ref 20]. The strength is scale, timeliness, and transparent construction of text measures with some human validation. The limitation is interpretability and endogeneity: “AI chatter” may reflect investor relations strategy rather than operational reality, and sentiment may be confounded by macro shocks.
2.21 Understanding Firms' AI Efforts and Their Economic Impact
This working paper synthesizes firm-level AI data sources and evidence on AI’s economic effects, arguing measurement is central and offering a framework for choosing among AI measures [Ref 21]. Its contribution is meta-scientific clarity: it explains why papers disagree and how to select measures aligned with specific questions. Its limitation (for your purposes) is that it is a review/synthesis; the opportunity lies in operationalizing its framework into new datasets and causal studies.
2.22 The effects of AI on firms and workers
This piece synthesizes evidence and describes a method to measure firm-level AI investments using job postings and resume data, emphasizing AI worker shares and reporting associated firm growth and workforce effects, while also noting occupational shifts toward educated/technical workers and rising concentration [Ref 22]. It provides a concrete measurement recipe and a set of empirically grounded hypotheses. As a non-academic-article format (based on the excerpt), it should be used cautiously as a pointer to underlying studies and measures.
2.23 GPTs are GPTs: An Early Look at the Labor Market Impact Potential of Large Language Models
This working paper proposes a rubric to assess occupational task exposure to LLM capabilities using human expertise and GPT-4 classifications, estimating large portions of tasks and jobs are exposed to LLMs and that tooling can substantially scale impacts [Ref 23]. Its strength is methodological novelty and policy relevance; its limitation is that it measures *technical exposure*, not realized adoption or equilibrium labor market outcomes.
2.24 Financial Fraud Detection Based on Machine Learning: A Systematic Literature Review
This systematic literature review uses the Kitchenham approach, selecting 93 articles and summarizing popular ML techniques (SVM, ANN), common fraud types (credit card fraud), and evaluation metrics, and it identifies issues and gaps [Ref 24]. It is strong methodologically as an SLR and offers a mature adjacent domain where AI impact can be measured through detection performance and operational deployment constraints. A gap remains in bridging “model performance” to organizational impact (loss reduction, deterrence, fairness, false positive costs).
2.25 Autonomous Vehicle Technology: A Guide for Policymakers
This policy guide highlights benefits and policy challenges of self-driving vehicles, including liability regulation and privacy/data control [Ref 25]. It is not directly connected to your finance/accounting focus but is useful as an analog: it shows that AI impacts are constrained by governance, liability, and data rights—issues that equally arise for AI in auditing, disclosures, and compliance.
2.26 The future of skills: employment in 2030
This report foregrounds the role of learning and skills in future employment. It is useful background for AI impact on skills and workforce transitions, but the provided excerpt does not supply specific empirical methods or results to leverage directly in high-impact academic work.
2.27 Polanyi’s Paradox and the Shape of Employment Growth
This working paper offers a conceptual and empirical overview of computerization’s effects on employment structure, emphasizing that routine codifiable tasks are most automatable while adaptability and common sense remain difficult; it highlights complementarities and cautions against overstating substitution [Ref 27]. It is foundational for AI-labor impact framing and helps interpret why AI adoption may reorganize tasks rather than eliminate jobs wholesale.
2.3 Comparative Summary Table
| Ref | Focus / Scope | Method or Approach | Key Contribution | Main Limitation or Gap |
|---|---|---|---|---|
| [Ref 1] | AI vs measured productivity (“productivity paradox”) | Conceptual framework; reviews evidence; proposes explanations | Implementation lags + intangible complementary capital framing | Limited direct micro causal identification in provided text |
| [Ref 2] | AI and earnings management in Nigerian SMEs | Quantitative survey; PLS-SEM; bootstrapping | Tests multiple AI-related constructs; reports strong associations | Potential common-method bias; extreme R²; limited generalizability |
| [Ref 3] | AI + blockchain impacts on accounting profession | Narrative comprehensive review | Maps AI tech types and applications; skills/governance concerns | Descriptive; lacks outcome-based causal evidence |
| [Ref 4] | Ethical challenges of AI in marketing for social good | Systematic ethical analysis | Multi-stakeholder tensions; beyond deontological principles | Limited operational metrics or causal tests |
| [Ref 5] | AI discourse and earnings management (China A-share, 2005–2024) | Text analysis of annual reports; fixed-effects panels | Distinguishes RM vs AM effects; scalable proxy | AI term frequency may be signaling/endogenous; adoption vs talk |
| [Ref 6] | AI and entrepreneurship | Conceptual analysis + research agenda | Task-level augmentation/replacement framing; liabilities highlighted | Needs empirical operationalization |
| [Ref 7] | AI in auditing: quality, fees, labor | Resume-based AI workforce measure + interviews | Evidence of improved audit quality and fee changes; lagged labor effects | Mechanisms and task-level pathways not fully isolated |
| [Ref 8] | LLMs vs analysts for core earnings (summary) | Describes LLM prompting workflows and evaluations | Shows structured prompts outperform naive use | Secondary source; must cite underlying study for publication |
| [Ref 9] | Earnings calls bibliometric review + AI prospects | Bibliometrics (14,437 articles) + snowball | Maps themes; motivates AI multi-modal analysis | Bibliometrics can’t validate measures/causality |
| [Ref 10] | Same as Ref 9 (duplicate) | Same as Ref 9 | Reinforces AI multi-modal opportunity | Duplicate |
| [Ref 11] | Patent law challenges from AI | Legal analysis with illustrative examples | Highlights institutional/legal constraints on AI commercialization | Not an impact evaluation; examples may not generalize |
| [Ref 12] | AI mentions, firm maturity, and performance (Finland services) | Keyword frequency in annual reports; ROA/ROE; OLC/TOE framing | Adds maturity contingency and adoption strategy framing | Proxy validity/endogeneity; thesis-level constraints |
| [Ref 13] | AI valuation in cross-border M&A | Random Forest Regressor vs EV/EBITDA benchmark | Claims improved valuation accuracy/robustness | Needs clearer ground truth and decision-use validation |
| [Ref 14] | AI, accounting information, earnings management (bibliometric) | Bibliometric (per abstract) | Signals consolidation of niche literature | Full text not provided; details unknown |
| [Ref 15] | “Intuitive intelligence” perspective | Narrative/speculative | General motivation | Not research-grade evidence |
| [Ref 16] | ML for climate change | Gap-focused agenda | Shows how to define high-impact AI problems | Different domain; indirect relevance |
| [Ref 17] | AI in earnings report processing (practitioner) | Descriptive claims; process narrative | Provides testable process-level hypotheses | Not validated academic evidence; potential marketing bias |
| [Ref 18] | LLM core earnings analysis (summary) | Summarizes “lazy” vs sequential prompts | Emphasizes prompt structure; scalability/cost claims | Secondary source; need underlying paper |
| [Ref 19] | AI/data analytics for disclosure compliance | Advisory framework | Identifies disclosure risk domains and AI uses | Needs translation into testable research designs |
| [Ref 20] | AI sentiment/uncertainty in earnings calls post-ChatGPT | Text-as-data; keyword + lexicon interactions; large corpus | Real-time measures of AI optimism vs uncertainty | Endogeneity; talk vs action; needs linkage to outcomes |
| [Ref 21] | Review of firm AI measures and economic effects | Measurement framework + synthesis | Explains why AI datasets differ and how to choose measures | Calls for new measurement/causal evidence |
| [Ref 22] | AI effects on firms/workers (synthesis) | Describes job postings + resume-based AI investment measure | Concrete measurement recipe and firm outcomes narrative | Primarily synthesis; underlying studies needed |
| [Ref 23] | LLM labor market exposure potential | Rubric + human/GPT-4 labeling of tasks | Quantifies exposure; tooling amplifies impact | Technical exposure ≠ realized adoption/equilibrium outcomes |
| [Ref 24] | ML fraud detection SLR | Kitchenham SLR; 93 articles | Maps methods/metrics; identifies gaps | Weak bridge from performance to org/economic impact |
| [Ref 25] | Autonomous vehicles policy issues | Policy analysis | Liability/privacy/data governance analogies | Domain mismatch to finance/accounting focus |
| [Ref 26] | Future skills and employment | Workforce/skills report | Highlights skills planning importance | Limited methods/results in provided excerpt |
| [Ref 27] | Polanyi’s paradox and employment growth | Conceptual + empirical labor framing | Complementarity vs substitution; routine task automation | Not AI-specific measurement; needs modern AI linkages |
2.4 Overall Assessment of the Provided Sources
Collectively, your sources cover (i) macro-level impact puzzles (productivity paradox), (ii) firm-level measurement and synthesis, (iii) accounting/auditing/earnings management applications, (iv) disclosure analytics via earnings calls, and (v) governance/ethics/legal constraints. This is a strong interdisciplinary foundation, especially for publishable work at the intersection of AI measurement, disclosure, and assurance.
What is missing is a cohesive empirical pipeline that links *specific AI investments* → *specific process changes* → *auditable intermediate outputs* → *financial reporting/audit quality outcomes* → *market and real outcomes*. Many sources rely on “AI mentions” or broad AI-worker measures; fewer directly observe AI tool deployment, task reallocation, model governance, and controls. That end-to-end chain is where high-impact contributions are currently most publishable.
3. The Broader Research Landscape
Research in this field has established several complementary traditions. First, an economics tradition treats AI as a general-purpose technology and studies diffusion, productivity, and labor-market task reallocation. Second, an accounting/finance tradition studies how AI changes information production (forecasting, disclosure), assurance (audit technology), and opportunistic reporting behavior (earnings management, fraud). Third, an information systems and management tradition examines technology adoption, organizational complements, and governance (capability building, sourcing, culture, and controls). Fourth, an ethics/law tradition analyzes accountability, transparency, privacy, IP/patents, and downstream harms or inequities.
Landmark lines of work include: task-based automation and routine-biased technological change; “text as data” methods in finance using corporate filings and calls; and machine learning in fraud detection and risk analytics. Generative AI/LLMs have intensified interest in measurement because they can both (a) change firm operations and (b) change *how researchers measure* firm behavior (automated extraction/classification). From a publication perspective, the most influential papers tend to either (i) introduce a validated new measure (dataset/method), or (ii) use compelling quasi-experimental identification to estimate causal effects, or (iii) provide a unifying framework that resolves contradictions across prior findings (measurement reconciliation, mechanisms, boundary conditions).
Ongoing debates include: whether AI primarily augments or displaces labor; whether impacts concentrate in superstar firms and increase market power; how quickly productivity gains diffuse beyond early adopters; and how to regulate AI without stifling innovation. A crucial unresolved question is attribution: when performance improves, is AI the driver or are better-governed firms simply more likely to adopt and to report adoption? Your sources explicitly underline this measurement and endogeneity challenge,,.
Key publication venues vary by angle. For economics and finance: leading journals in economics, finance, and accounting; for auditing and disclosure: top accounting journals; for interdisciplinary technology and management: top management and information systems journals; and for methods-heavy work: outlets that welcome data and measurement contributions. (Specific venue targeting is outlined in Section 8.)
4. Strengths of Existing Research
A major strength is the increasingly sophisticated measurement toolkit. The literature is moving beyond simplistic “AI = patents” proxies toward richer firm-level indicators such as AI workforce composition derived from resume/job posting data and detailed disclosure text measures [Ref 7], [Ref 20], as well as explicit frameworks for choosing among measures depending on whether you study invention, adoption, outsourcing, or perceptions [Ref 21]. This improves interpretability and helps align empirical design with theory.
Second, in auditing, evidence has progressed from conjecture to measurable outcomes. The association between AI investments and reduced restatements, plus fee changes and delayed labor displacement, provides a concrete foundation for mechanism testing and replication [Ref 7]. The use of interviews to support empirical interpretation is an additional strength because it reduces the risk that AI measures are purely statistical artifacts.
Third, the disclosure and earnings-management literature is beginning to separate channels and constructs. Distinguishing real activities manipulation from accrual-based earnings management, and finding differential relationships with AI-related discourse, is exactly the kind of nuance reviewers look for because it aligns with theory about constraints, detectability, and auditability [Ref 5]. Similarly, the earnings-calls text-as-data approach provides high-frequency, scalable signals of managerial attention, sentiment, and uncertainty that can be linked to subsequent decisions and performance [Ref 20].
5. Weaknesses and Limitations of Existing Research
The dominant limitation is construct validity and endogeneity in “AI adoption” proxies. Keyword frequency in annual reports or earnings calls can capture strategic signaling, marketing, or managerial optimism rather than actual process-level implementation [Ref 5], [Ref 12], [Ref 20]. Even AI workforce measures can reflect centralization strategies or reporting differences rather than effective deployment. Babina explicitly warns that different datasets capture different underlying objects and can therefore lead to different conclusions [Ref 21]. For publication, you must show why your chosen AI measure is the right one for your specific causal claim.
A second limitation is causal identification. Many studies are correlational (fixed effects help but do not eliminate time-varying confounding), or they rely on cross-sectional survey perceptions (high risk of common-method bias). For example, the PLS-SEM SME study reports extremely high R² (0.990), which can be a red flag for reviewers because it suggests overlapping constructs, response bias, or model overfit rather than true explanatory power [Ref 2]. High-impact journals will demand robustness, alternative explanations, and ideally quasi-experimental variation.
Third, mechanisms are often asserted rather than tested. Macro narratives emphasize implementation lags and intangible complements [Ref 1], but micro evidence rarely measures the complements directly (data pipeline maturity, governance controls, training, process redesign). Similarly, audit evidence shows outcomes improve with AI hiring, but specific task-level mechanisms (risk assessment, anomaly detection, substantive testing automation, client selection) need sharper measurement [Ref 7].
Fourth, governance, ethics, and law are frequently treated as “discussion sections” instead of measurable moderators. Yet ethical controversies and stakeholder tensions can shape adoption choices, customer trust, regulatory enforcement, and litigation risk [Ref 4], [Ref 19], [Ref 11]. Treating governance as measurable (e.g., board oversight structures, AI policy disclosures, audit committee expertise, model risk controls) is a major opportunity.
6. Critical Research Gaps
1) Validated, multi-dimensional firm AI measurement that separates “talk” from “use” Unknown: How to reliably distinguish AI implementation from AI-related disclosure strategy across firms and time. Why it matters: Without this, estimates of AI’s impact on productivity, reporting quality, or labor are biased and hard to reconcile across studies [Ref 21]. Evidence it’s open: Multiple sources rely on text proxies and acknowledge measurement centrality; divergence across datasets remains unresolved [Ref 21], [Ref 5], [Ref 20]. Impact: HIGH.
2) Causal effects of AI adoption on earnings management and reporting quality (beyond associations) Unknown: Whether AI actually constrains or shifts earnings management, and through which channels (auditability, transparency, internal controls). Why it matters: This affects investor protection, capital allocation, and regulatory design. Evidence it’s open: Mixed findings across RM vs AM and across contexts, with text proxies that can be endogenous [Ref 5]; SME perceptual results may not generalize [Ref 2]. Impact: HIGH.
3) Mechanism-level understanding of how AI improves audit quality and how long labor impacts take to materialize Unknown: Which AI tasks/tools deliver restatement reductions and fee changes, and how centralization shapes outcomes. Why it matters: Helps firms and regulators target effective AI governance and avoid harmful displacement or quality trade-offs. Evidence it’s open: Outcomes are documented, but mechanisms are mostly inferred; interviews indicate central development and broad use, but task pathways remain under-measured [Ref 7]. Impact: HIGH.
4) Multi-modal disclosure analytics (text + audio + video) for earnings calls/videocasts Unknown: Whether vocal tone, facial cues, and visual presentation add incremental predictive power over text for risk, future performance, and misreporting. Why it matters: Corporate disclosure is evolving to richer media; researchers risk being left behind if they only analyze transcripts [Ref 10]. Evidence it’s open: Bibliometric review explicitly calls for AI-enabled voice/image analysis and highlights videocasts as a frontier [Ref 10]. Impact: MEDIUM-HIGH.
5) AI governance, ethics, and compliance as measurable moderators of AI’s economic impact Unknown: Which governance practices (oversight, documentation, audit trails, fairness testing, disclosure controls) enable benefits while reducing harms. Why it matters: Ethical controversies and compliance risks can dominate realized value, especially in marketing and public disclosures [Ref 4], [Ref 19]. Evidence it’s open: Ethical tensions are analyzed conceptually, but empirical moderator tests remain sparse [Ref 4]. Impact: MEDIUM-HIGH.
6) Bridging model performance to organizational/economic impact in fraud detection and compliance Unknown: When improved detection metrics translate into loss reduction, deterrence, and net value after false positives and operational constraints. Why it matters: Many AI deployments “work” in AUC terms but fail operationally. Evidence it’s open: SLR catalogs techniques and gaps but highlights unresolved issues and limitations [Ref 24]. Impact: MEDIUM.
7. Recommended Research Directions
7.1 Priority Research Opportunities
Recommendation 1: “AI Talk vs AI Walk” — A validated adoption index for finance/accounting contexts Research Question: Can we build a composite firm-year AI adoption measure that distinguishes actual capability-building and deployment from disclosure signaling, and does it predict real operational changes? Why It Matters: Resolves the core measurement problem highlighted in firm-level AI synthesis [Ref 21] and reduces bias in downstream impact studies. Addresses Gap: Gap 1. Suggested Methodology:
- Construct multiple measures: (i) AI keyword frequency in annual reports (as in [Ref 5], [Ref 12]); (ii) AI chatter/sentiment in earnings calls (as in [Ref 20]); (iii) AI workforce intensity (conceptually aligned with [Ref 7] and measurement discussion in [Ref 22]); (iv) evidence of centralized AI teams or outsourcing signals where observable.
- Validate against external benchmarks: e.g., observed AI product launches, process automation announcements, or documented AI governance structures (from public disclosures).
- Use factor models or latent variable modeling, but include strong out-of-sample validation to avoid the “R²=0.99” credibility problem [Ref 2].
Expected Contribution: A publishable dataset/method paper + a reusable measurement index for the field. Publication Potential: High in journals receptive to measurement innovation in accounting, finance, or information systems; strong chance if released as a replicable dataset with transparent code.
Recommendation 2: Causal impact of AI adoption on Real vs Accrual earnings management using quasi-experiments Research Question: Does AI adoption causally reduce RM and/or AM, or does it shift manipulation from one type to another? Why It Matters: Builds directly on the RM vs AM distinction and mixed evidence in [Ref 5], and tests the claim that AI increases transparency and auditability. Addresses Gap: Gap 2. Suggested Methodology:
- Use difference-in-differences designs around plausibly exogenous adoption shocks: e.g., abrupt rollout of AI audit tools at certain audit firms (linked to [Ref 7]) or policy/regulatory changes that alter AI feasibility/cost.
- Combine firm fixed effects with event-study plots, placebo tests, and sensitivity analyses.
- Measure RM and AM with standard accounting models; ensure robustness across specifications.
Expected Contribution: Causal evidence on whether AI constrains opportunism or changes its form. Publication Potential: Strong for top accounting/finance outlets if identification is credible and mechanisms are tested.
Recommendation 3: Mechanism mapping in AI-enabled auditing (task-level decomposition + governance) Research Question: Which specific audit tasks and organizational arrangements (centralized AI teams, tool governance, training) mediate the observed quality gains from AI investments? Why It Matters: [Ref 7] shows outcomes; your opportunity is to explain *how* and *under what conditions* they arise. Addresses Gap: Gap 3 and Gap 5. Suggested Methodology:
- Mixed-method design: replicate/extend resume-based AI workforce measures [Ref 7] and add audit-task proxies (e.g., changes in audit report timing, scope indicators, restatement types, internal control weaknesses).
- Collect qualitative data (structured interviews, or coded partner commentary where available) to classify AI use cases.
- Test mediation: AI workforce → audit process proxies → restatements/fees.
Expected Contribution: Mechanistic evidence that is highly valued by reviewers and practitioners. Publication Potential: High in auditing-focused journals; also attractive to regulators.
Recommendation 4: Multi-modal earnings call analytics for AI uncertainty, credibility, and future misreporting risk Research Question: Do vocal and visual cues in earnings call videocasts add incremental predictive power for AI-related uncertainty, risk disclosure, and subsequent reporting quality beyond transcript text? Why It Matters: Earnings call research is evolving to videocasts; bibliometric work explicitly flags voice/image recognition as a frontier [Ref 10]. Addresses Gap: Gap 4. Suggested Methodology:
- Build a dataset of videocast earnings calls (where available) and align with transcripts.
- Extract: text sentiment/uncertainty (as in ); audio prosody (pitch variability, speech rate); visual cues (facial action units, gaze, presentation dynamics).
- Predict outcomes: forecast errors, volatility, subsequent restatements, or changes in guidance.
Expected Contribution: A novel disclosure measurement contribution; strong “new data + new method” appeal. Publication Potential: High in information systems, accounting information systems, or finance outlets that value alternative data.
Recommendation 5: LLM governance in financial reporting workflows—prompting, audit trails, and error taxonomy Research Question: What governance and workflow designs make LLM-based extraction/classification in financial reporting reliable, and what error modes remain systematic? Why It Matters: LLM performance depends strongly on structure (“lazy” vs sequential prompts) [Ref 18]; firms are adopting AI for disclosure processing and compliance [Ref 17], [Ref 19]. Addresses Gap: Gap 5 and partly Gap 1. Suggested Methodology:
- Controlled experiments using real 10-K sections: compare prompting strategies, retrieval-augmented setups (if used), and structured step decomposition similar to sequential prompting described in [Ref 8], [Ref 18].
- Create an accounting-specific error taxonomy (concept confusion like EBITDA vs core earnings; recurring vs non-recurring misclassification).
- Evaluate reproducibility, cost, and auditability (traceable citations to source text).
Expected Contribution: Practical and publishable evidence on reliable GenAI use in regulated reporting contexts. Publication Potential: Strong in accounting information systems, auditing tech, or computational finance venues.
Recommendation 6: AI ethics and “social good” claims as measurable determinants of trust and economic outcomes Research Question: When firms claim AI-for-social-good positioning, does it change customer/investor trust, regulatory scrutiny, and long-run performance—and how do ethical tensions moderate these effects? Why It Matters: Ethical controversies are central to AI in stakeholder-facing functions [Ref 4]; disclosure risk management is increasingly strategic [Ref 19]. Addresses Gap: Gap 5. Suggested Methodology:
- Text-as-data classification of AI ethics and social-good language in sustainability and annual reports; combine with event studies around controversies or enforcement actions.
- Multi-stakeholder outcomes: customer churn proxies, litigation/activism incidence, cost of capital proxies.
Expected Contribution: Moves ethics from “normative discussion” to measurable economic mechanisms. Publication Potential: High in business ethics and interdisciplinary management journals if measurement is rigorous.
7.2 Quick-Win Opportunities
1) Replication + robustness note on AI chatter measures post-ChatGPT Use the method described for counting AI-related sentences and sentiment/risk interactions in earnings calls [Ref 20] and test robustness to alternative AI keyword sets and lexicons, plus simple predictive validity (e.g., relation to subsequent capex or R&D intensity). Publishable as a short empirical note if you add a novel validation angle.
2) Prompting experiment on a narrow accounting task (6–12 months) Using the structured vs minimal prompting contrast emphasized in LLM core earnings summaries [Ref 18], test on one well-defined classification task (non-recurring items, segment extraction). Output: benchmark dataset + error analysis + governance recommendations.
3) Systematic mapping review: “AI and earnings management” measurement typology Build a small systematic review that categorizes how studies measure “AI adoption” (mentions, patents, workforce, vendor tools) and how they measure earnings management (RM, AM, restatements), explicitly following the measurement centrality argument [Ref 21]. This can be publishable in a review-friendly outlet and sets up your empirical agenda.
8. How to Maximise Your Publication Success
Target journals and conferences based on *your core contribution type*. If you produce a new measure/dataset (validated AI adoption index; multi-modal earnings calls dataset), journals that value data and methods contributions are more receptive than purely theory outlets. If you produce causal estimates (DiD/event studies on AI adoption shocks), aim at top accounting, finance, and economics outlets where identification is paramount. If you produce governance/ethics-mechanism work, business ethics and interdisciplinary management outlets will be more receptive—provided you quantify constructs, not just argue them.
Reviewers in this field typically look for: (i) a defensible AI measure (what exactly is being captured?), (ii) a clear theoretical mechanism (why should AI increase/decrease manipulation, quality, productivity?), (iii) credible identification or strong validation (out-of-sample tests, falsification tests), and (iv) interpretability (not just predictive accuracy). The biggest avoidable rejection reason is overstating what your proxy measures. If you use AI keyword frequency, explicitly frame it as “AI-related disclosure” and test when it aligns with observable capability-building. Use Babina’s measurement framework as your conceptual guardrail: always state whether you measure invention, adoption, internal build, outsourcing, or perceptions [Ref 21].
Common pitfalls include: (a) treating LLM outputs as ground truth without auditing error modes, (b) ignoring implementation lags (benefits may appear years later, consistent with lag arguments [Ref 1] and lagged labor effects in audit [Ref 7]), (c) failing to separate RM and AM (a strength you can borrow from [Ref 5]), and (d) ignoring governance/ethics as moderators even though they plausibly determine success or backlash [Ref 4], [Ref 19].
Collaboration strategy: pair domain expertise (accounting/audit/finance) with ML/NLP expertise and, if doing multi-modal work, with speech and computer vision expertise. For causal work, collaborate with an applied econometrician who can stress-test identification. For publication positioning, explicitly connect to the productivity paradox narrative: even if your outcome is audit quality or earnings management, frame it as “micro evidence on the complements and lags of AI impact,” which resonates with macro audiences [Ref 1].
9. Suggested Research Framework
A coherent agenda can be sequenced in three stages. Stage 1 (0–12 months): solve measurement and reliability foundations. Build and validate a multi-dimensional AI adoption index that separates talk from use (Section 7.1, Rec 1), and run a narrow prompting governance experiment for LLM accounting tasks (Rec 5 quick-win). This produces early publications (dataset/method + short empirical note) and creates infrastructure for later causal work.
Stage 2 (12–24 months): deploy the validated measures into causal designs. Use quasi-experimental variation to estimate AI’s causal impact on RM vs AM, restatements, and disclosure quality (Rec 2), and deepen auditing mechanisms (Rec 3). Stage 3 (24+ months): expand into frontier modalities and governance: multi-modal videocast analysis (Rec 4) and ethics/governance moderators (Rec 6). Across all stages, keep the macro motivation explicit: AI’s impact is mediated by intangible complements, governance, and diffusion lags—your micro evidence can help resolve the broader productivity paradox [Ref 1].
10. Conclusion
The highest-priority gap is credible measurement: many apparent “AI effects” may be effects of disclosure strategy, investor sentiment, or pre-existing governance quality rather than AI deployment itself. This is why measurement centrality is repeatedly emphasized in firm-level AI synthesis and why keyword-based proxies must be validated rather than assumed.
The most actionable high-impact recommendation is to build a validated, multi-dimensional firm AI adoption index (separating talk from use) and then use it in a quasi-experimental design to estimate AI’s causal effects on real vs accrual earnings management and reporting quality. This combination—new measure + credible causal inference—maximizes your chance of publication in high-impact journals while producing knowledge that is genuinely useful to firms, auditors, and regulators.
11. References (IEEE format)
[1] E. Brynjolfsson, D. Rock, and C. Syverson, "Artificial Intelligence and the Modern Productivity Paradox: A Clash of Expectations and Statistics," National Bureau of Economic Research, 2017, doi: 10.3386/w24001.
[2] "Advances in Artificial Intelligence and Machine Learning; Research 6 (2) 5142-5159 Received 24-12-2025; Accepted 01-03-2026; Published online 07-03-2026," n.d. [Online]. Available: https://www.oajaiml.com/uploads/archivepdf/583562285.pdf.
[3] Y. Zhang, F. Xiong, Y. Xie, X. Fan, and G. Hai-feng, "The Impact of Artificial Intelligence and Blockchain on the Accounting Profession," IEEE Access, 2020, doi: 10.1109/access.2020.3000505.
[4] E. Hermann, "Leveraging Artificial Intelligence in Marketing for Social Good—An Ethical Perspective," Journal of Business Ethics, 2021, doi: 10.1007/s10551-021-04843-y.
[5] "Research On the Impact of Artificial Intelligence on Corporate Earning Management," n.d. [Online]. Available: https://www.shs-conferences.org/articles/shsconf/pdf/2025/16/shsconf_icfmde2025_03015.pdf.
[6] D. Chalmers, N. MacKenzie, and S. Carter, "Artificial Intelligence and Entrepreneurship: Implications for Venture Creation in the Fourth Industrial Revolution," Entrepreneurship Theory and Practice, 2020, doi: 10.1177/1042258720934581.
[7] A. Fedyk, J. Hodson, N. V. Khimich, and T. Fedyk, "Is artificial intelligence improving the audit process?," Review of Accounting Studies, 2022, doi: 10.1007/s11142-022-09697-x.
[8] "Forecasting Earnings, Artificial Intelligence (AI) Versus Equity Analysts," n.d. [Online]. Available: https://www.interactivebrokers.com/campus/ibkr-quant-news/….
[9] J. M. Bravo and R. Maia, "Twenty-five years (1998-2023) of Earnings Disclosure with Conference Calls: a Bibliometric Review focusing on Artificial Intelligence," AIS Electronic Library (AISeL), 2024. [Online]. Available: https://core.ac.uk/download/639871997.pdf.
[10] J. M. Bravo and R. Maia, "Twenty-five years (1998-2023) of Earnings Disclosure with Conference Calls: a Bibliometric Review focusing on Artificial Intelligence," AISEL, 2024, doi: 10.54499/uidb/04152/2020.
[11] M. Hashiguchi, "The Global Artificial Intelligence Revolution Challenges Patent Eligibility Laws," DigitalCommons@UM Carey Law, 2017. [Online]. Available: https://core.ac.uk/download/153518163.pdf.
[12] F. Lolo, "Firm Maturity, AI Adoption, and Financial Performance: A Study of Publicly Listed," 2025. [Online]. Available: https://core.ac.uk/download/652900364.pdf.
[13] S. Petro-Korhonen El Bouchtili, "Saana Petro-Korhonen El Bouchtili AI-DRIVEN VALUATION TECH-NIQUES IN CROSS-BORDER MER-GERS AND ACQUISITIONS Enhancing Accuracy in Emerging Markets," 2025. [Online]. Available: https://core.ac.uk/download/657107044.pdf.
[14] "The Impact of Artificial Intelligence on Accounting Information and Earnings Management: Bibliometric Analysis," n.d. [Online]. Available: https://ideas.repec.org/a/gam/jjrfmx/v19y2026i1p90-d1846041.html.
[15] Y. Dzbanovskii, L. Korotenko, and N. Nechai, "Section 03. Smart Solutions in IT," Видавництво НГУ, 2017. [Online]. Available: https://core.ac.uk/download/132414646.pdf.
[16] L. H. Kaack et al., "OPUS 4 | Tackling Climate Change with Machine Learning," OPUS 4 (Zuse Institute Berlin), 2022, doi: 10.1145/3485128.
[17] "How AI is Transforming Earnings Report Processing in 2025 - Daloopa," n.d. [Online]. Available: https://daloopa.com/blog/analyst-best-practices/….
[18] "AI-Powered Core Earnings Analysis: A New Frontier in Financial Reporting | Harvard Business School AI Institute," n.d. [Online]. Available: https://d3.harvard.edu/ai-powered-core-earnings-analysis-a-new-frontier-in-financial-repor….
[19] "Using Data Analytics and Artificial Intelligence for Public Disclosures," n.d. [Online]. Available: https://corpgov.law.harvard.edu/2024/02/….
[20] "AI Optimism and Uncertainty: What Can Earnings Calls Tell Us Post-ChatGPT?," n.d. [Online]. Available: https://www.stlouisfed.org/on-the-economy/2024/….
[21] "NBER WORKING PAPER SERIES," n.d. [Online]. Available: https://www.nber.org/system/files/working_papers/w35123/w35123.pdf.
[22] "The effects of AI on firms and workers | Brookings," n.d. [Online]. Available: https://www.brookings.edu/articles/the-effects-of-ai-on-firms-and-workers/.
[23] T. Eloundou, S. Manning, P. Mishkin, and D. L. Rock, "WORKING PAPER," arXiv (Cornell University), 2023, doi: 10.48550/arxiv.2303.10130.
[24] A. Ali et al., "Financial Fraud Detection Based on Machine Learning: A Systematic Literature Review," Applied Sciences, 2022, doi: 10.3390/app12199637.
[25] J. Anderson, N. Kalra, K. Stanley, P. Sørensen, C. Samaras, and O. Oluwatola, "Autonomous Vehicle Technology: A Guide for Policymakers," RAND Corporation eBooks, 2016, doi: 10.7249/rr443-2.
[26] H. Bakhshi, J. M. Downing, M. A. Osborne, and P. Schneider, "The future of skills: employment in 2030," Oxford University Research Archive (ORA) (University of Oxford), 2017. [Online]. Available: https://ora.ox.ac.uk/objects/uuid:86577437-1353-4743-8520-401c1f99ad1b.
[27] D. Autor, "Polanyi’s Paradox and the Shape of Employment Growth," National Bureau of Economic Research, 2014, doi: 10.3386/w20485.
12. Key Literature From Broader Field (TIER 3 — Verify Before Citing)
- [TRAINING — verify before citing] E. Brynjolfsson and A. McAfee, *The Second Machine Age*, 2014.
- [TRAINING — verify before citing] R. Solow, “You can see the computer age everywhere but in the productivity statistics,” 1987 (often referenced as the “Solow paradox”).
- [TRAINING — verify before citing] T. A. Hassan, S. Hollander, A. Kalyani, L. van Lent, M. Schwedeler, and A. Tahoun, “Economic Surveillance Using Corporate Text,” 2024 (mentioned in as a working paper).
- [TRAINING — verify before citing] T. Loughran and B. McDonald, financial sentiment lexicon / dictionary (used in [Ref 20] description).
- [TRAINING — verify before citing] B. Kitchenham, guidelines/method for systematic literature reviews in software engineering (the “Kitchenham approach” referenced in [Ref 24]).
- [TRAINING — verify before citing] J. Tobin, *q* theory / market valuation linkage to intangible investment (relevant to intangible capital discussion; include only if you explicitly use it).
- [TRAINING — verify before citing] Standard RM/AM measurement models in earnings management (e.g., common accrual models and real activities manipulation models)—cite the specific canonical papers once you select them.
- [TRAINING — verify before citing] Retrieval-Augmented Generation (RAG) and LLM evaluation/governance frameworks (if you use them, cite specific standards/papers you adopt).
أنشئ تقريراً كهذا عن موضوعك.
ابدأ بحثك المجاني