The debate on synthetic data held at Stevens Institute of Technology, as part of the Insights Association’s AI Ignite Conference in October 2024, brought to the forefront a growing discourse on whether synthetic data can fully replace human-generated insights. Panelists included Generation1.ca Founder and CEO Arundati Dandapani, Dwayne Allen, CTO of Unisys, Jess Hall, Principal at Hallway Studios, and John Bremer, CEO of Phantom5 Solutions. Each of these experts offered diverse perspectives on this critical topic.

As the opposition leader, Dandapani, joined by Bremer, argued that while synthetic data offers efficiency and scalability, it cannot replace the richness of human-generated insights, particularly when it comes to capturing the nuanced and evolving experiences of multicultural and multilingual communities. She emphasized the need for human oversight and robust data and AI governance, ensuring that synthetic data serves as a complement rather than a replacement to human intelligence, especially in contexts where data scarcity cannot be solved by more data scarcity. Moreover, Dandapani pointed out the challenges around the varied definitions of synthetic data, noting that it has been in use in government contexts, like the US Census Bureau, since the 1970s, although often under different nomenclatures. Establishing a common definition is critical for effective usage and regulation.
In contrast, proponents Hall and Allen focused on synthetic data’s transformative potential in sectors like healthcare and finance, where privacy concerns restrict the use of real data. They argued that synthetic data could democratize access to innovation, enabling organizations to develop AI models without compromising data privacy. Allen cited examples from the financial services and healthcare sectors, where synthetic data facilitated accelerated innovation during the COVID-19 pandemic, particularly in generating safe, scalable datasets for drug development.
This debate paralleled discussions from a December 2023 Town Hall organized by the Insights Association, where experts like Benjamin de Seingalt, Chris Robson, Damion Taylor, Katie Gross, and Zachary Richard examined the benefits and challenges of synthetic data. Their dialogue had similarly centered on the ethical and legal implications of AI-generated data, especially the risks associated with blending deterministic and probabilistic data in the presence of biases and the “black box” nature of AI solutions. The panelists stressed the need for clear guidelines and transparency to ensure stakeholders are aware of the synthetic nature of the data they utilize, especially as open-source models gain traction. The consensus was that while synthetic data can simulate realistic data patterns, it may amplify existing biases from training datasets, making human oversight essential to maintain data integrity.

Both the Stevens Institute debate and the Insights Association’s Town Hall underscored a growing consensus: while synthetic data holds significant potential for innovation, it should not—and indeed cannot—replace human-generated insights. The need for industry standards, regulatory clarity, and ethical guidelines remains paramount to navigate this complex landscape responsibly. As synthetic data evolves, businesses and organizations must balance its benefits with careful governance to ensure that data-driven decisions are inclusive and accurate, driving both business and social impact for stakeholders, customers and industry ecosystems including talent pipelines.

These debates culminate in a broader reflection on the pitfalls and potential of an extraordinary imagination, beyond realities and promise. Like good fiction and non-fiction, both human-generated and synthetic data serve as embodiments of truth—but at scale. Just as literature seeks to convey deeper truths through storytelling, synthetic data aims to unlock both quantitative and qualitative insights, pushing the boundaries of what we know. However, while the mathematical precision of synthetic data excels in structured, quantitative analysis, it currently arguably falls short in capturing nuanced, experience-rich perspectives that rely on human intuition and cultural, geospatial or perceptual understanding.

In her ESOMAR Congress 2024 paper, Rediscovering the Immigrant Journey: Mapping the Potential of Synthetic Data in Global Immigrant Stories, Dandapani explored the guidelines for leveraging synthetic data while emphasizing the human truths essential for understanding immigrant attitudes, needs, and values. Her research and works spanning her publications and presentations from Chicago to Athens (where she also won ESOMAR’s Best Paper of the Year Award in 2024), highlights the necessity of harmonizing both human and synthetic data to uncover deeper insights. The future lies not in choosing between these data types but in integrating them to achieve scalable yet profoundly human data-driven decisions.

1 Comment