EP 2: Data & Actuarial Science

2022-08-09 · Show: Data Science With Sam · 2729s · Source

Data, Risk, and Actuarial Science in Insurance

概览

This episode discusses how actuarial science uses data to quantify insurance risk, with Mary Pat Campbell explaining that the field is not only about mathematics but also about understanding real-world behavior, business processes, regulation, and data limitations.

A central theme is that actuarial work depends on knowing what data actually represents. The conversation moves from mortality tables and life annuities to COVID mortality, fraud detection, reporting lags, missing values, coding errors, and the professional standards actuaries use to maintain data quality.

The later discussion focuses on AI, machine learning, and data science in insurance. Mary Pat argues that these tools can be useful, but insurance models operate under regulatory and business constraints, so data scientists and actuaries need to work together with strong subject-matter understanding.

分段落总结

[00:03] Introduction to the Guest and Topic

[事实] The host introduces the episode as a coffee chat about actuarial science and the role of data in the field. [事实] Mary Pat Campbell is introduced as a life actuary working in insurance research, a fellow of the Society of Actuaries, and a member of the American Academy of Actuaries. [事实] The host notes her background in physics and applied math, her involvement with SOA modeling work, and her writing on mortality and public finance. [推测] The opening positions the episode for viewers who work at the intersection of data science, insurance, and actuarial practice.

[02:06] Actuarial Science as Risk Quantification

[事实] Mary Pat defines actuarial science as centered on quantifying risk, not only financial engineering. [事实] She explains that early actuarial work was closely tied to mortality tables and life annuities. [事实] She mentions Edmund Halley as an early figure involved in mortality-table development. [事实] She says actuarial work uses a broad “tool bag” of techniques from different areas to price risks in insurance, pensions, and other promise-making businesses. [推测] Her framing emphasizes actuarial science as applied, business-facing quantitative work rather than a purely academic discipline.

[05:01] Life Insurance, P&C Insurance, and Data Volume

[事实] Mary Pat says her specialty is life insurance, where mortality and morbidity are central. [事实] She contrasts life insurance and annuities, which can last decades, with property and casualty policies that often renew annually or even every six months. [事实] She says life-side mortality data often needs to be shared or aggregated across insurers because deaths are relatively infrequent and take time to observe. [事实] She notes that the SOA aggregates mortality data across the industry, including life insurance and pensions. [推测] The key data challenge differs by insurance line: life work needs long-horizon aggregation, while P&C can often use faster feedback loops.

[07:45] Using Past Data to Project Future Risk

[事实] Mary Pat says actuarial work must project into the future even though incoming data reflects the past. [事实] She discusses COVID mortality and says the mortality experience of 2020 and 2021 should not automatically become the future pricing assumption for life insurance. [事实] She compares this with the Spanish flu, saying period life expectancy dropped sharply in one year and then largely recovered the next year. [事实] She says reinsurance and retrocession exist to help handle unusually bad mortality years. [推测] The discussion warns against mechanically treating an extreme recent event as a permanent future baseline.

[09:55] Fraud, Underwriting, and Selection Risk

[事实] Mary Pat says fraud is less common on the life side than in property and casualty, where claims data can reveal suspicious patterns. [事实] She says that in life reinsurance, very large claims were investigated. [事实] She explains that life insurance fraud is often more likely to appear in underwriting than in faking a death. [事实] She describes asymmetric information or selection risk as cases where the insured party knows they carry more risk than the insurer recognizes. [推测] This section shows why actuarial judgment requires both quantitative analysis and practical knowledge of how people and businesses behave.

[12:16] Real Data, Fake Data, and Detecting Anomalies

[事实] Mary Pat says actuarial work requires knowing whether the data being examined is real or fake. [事实] She refers to a published research case involving Dan Ariely where disclosed data later showed strange patterns. [事实] She says the results did not appear suspicious until people examined the underlying data. [事实] She also mentions statisticians finding unnatural voting patterns in Russian election data. [推测] The shared point across these examples is that data validation can expose problems that are invisible from conclusions alone.

[15:52] Sponsor Break

[事实] The episode includes a sponsor segment for Eureka Ergonomic office chairs. [事实] The ad emphasizes home-office aesthetics, leather finishes, and furniture-inspired design.

[16:52] Actuarial Standards of Practice

[事实] The host asks about actuarial standards of practice, especially data quality and integrity. [事实] Mary Pat says credentialed actuaries in the United States must follow applicable ASOPs. [事实] She identifies ASOP 23 on data quality, ASOP 41 on actuarial communications, and ASOP 56 on modeling as especially important. [事实] She recommends that non-actuaries working with data also read these standards because they are short and high level. [推测] She treats the ASOPs as useful professional discipline, not only as formal compliance documents.

[18:58] COVID Data and Reporting Lags

[事实] Mary Pat says ASOP 23 does not require a full audit of every dataset, but it does require actuaries to understand how data can go wrong. [事实] She uses COVID mortality reporting to explain the difference between when deaths occur and when deaths are reported. [事实] She says apparent weekend drops or holiday drops in death counts can reflect reporting delays rather than real decreases in mortality. [事实] She connects this to insurance concepts such as incurred but not reported claims. [推测] The broader lesson is that time stamps and reporting processes can materially change how data should be interpreted.

[22:20] Correlation, Causation, and Data Field Meaning

[事实] The host raises the point that correlation does not imply causation. [事实] Mary Pat says actuarial science generally receives observational data from the world rather than experimental data. [事实] She says actuaries need to understand database fields and flags before using them. [事实] She gives an example of a policy-file flag being misinterpreted by a valuation company, requiring correction. [事实] She says actuarial malpractice lawsuits can arise from data-coding misunderstandings. [推测] This section makes metadata and field definitions as important as the numerical values themselves.

[24:53] Bad Data in Health and Financial Records

[事实] Mary Pat says automated underwriting and electronic health records can contain bad data. [事实] She gives an example of her recorded weight appearing as zero because she was not weighed during a doctor visit. [事实] She says analysts must distinguish a true zero from a missing-data value. [事实] She says the same issue appears in financial data when deciding whether a value is truly zero or simply not reported yet. [推测] Automated systems can amplify errors if analysts treat every recorded value as clean and meaningful.

[25:36] Medical Codes, Prescriptions, and Practical Knowledge

[事实] Mary Pat says billing or medical codes may reflect testing for a condition rather than an actual diagnosis. [事实] She gives an example of an old prescription remaining in her records even though she had not taken it for years. [事实] She says much of this knowledge is learned on the job and may be specific to the type of insurance being analyzed. [事实] She describes ASOPs as expected practices built from practical actuarial experience. [推测] Domain context is necessary because the same data field can mean different things depending on workflow and source system.

[27:27] Reasonability Checks, Units, and Currency

[事实] Mary Pat says ASOP 23 expects reasonability checks rather than superhuman investigation. [事实] She says analysts should understand plausible ranges for values such as blood pressure, glucose, or temperature. [事实] She gives examples of impossible values, such as blood pressure recorded as 1800 over 50 or temperature recorded in the wrong unit. [事实] She also mentions financial data errors involving currencies such as U.S. dollars versus Japanese yen. [推测] The practical habit she recommends is systematic checking of ranges, units, definitions, and scale before modeling.

[30:21] ASOP 23 for Non-Actuaries

[事实] The host says non-actuaries working with actuaries should understand basic ASOP guidance when doing data management or analytics for actuarial work. [事实] Mary Pat says actuaries appreciate when collaborators understand these expectations because actuaries must still perform the necessary checks. [事实] She says ASOP 23 is the main standard for data quality and is short, with much of it focused on definitions. [推测] Shared standards can make collaboration between actuaries, data analysts, and compliance teams more efficient.

[31:02] AI, Machine Learning, and the Future of Actuarial Work

[事实] The host asks whether AI and machine learning can help actuaries manage data quality, data integrity, and modeling. [事实] Mary Pat says she wants to avoid a utopian view and describe how things actually work. [事实] She looks back to the founding of the Casualty Actuarial Society and says casualty actuarial work historically had ties to statisticians. [事实] She says data science has strong statistical foundations. [推测] Her answer suggests that the relationship between actuarial work and data science is not new, even if the tools and terminology have changed.

[33:45] Collaboration Instead of Professional Barriers

[事实] Mary Pat says actuaries have periodically worried about outsiders taking their jobs, including statisticians, data mining specialists, GLM practitioners, and now AI or machine learning practitioners. [事实] She says she does not think building barriers is a good idea. [事实] She says actuaries are not expected to be the statistical or IT experts, but they need to understand enough about systems and models to use them properly. [事实] She says actuaries may use vendor software but often need to adjust it for actuarial purposes. [推测] The preferred model is a team structure where actuarial, statistical, data-science, and IT expertise are combined.

[35:40] Insurance Models Under Regulation

[事实] Mary Pat says insurance differs from companies like Amazon or Google because insurance is heavily regulated on multiple levels. [事实] She says insurance models cannot be fully unconstrained. [事实] She gives credit scoring in personal auto insurance as an example of a strong real correlation that regulators may dislike for pricing. [事实] She mentions a Washington state dispute over limiting the use of credit scoring. [推测] A technically predictive variable may still be unusable if it conflicts with regulatory, legal, or fairness constraints.

[37:04] How Data Scientists Can Help Insurance

[事实] Mary Pat says data scientists could help by developing approaches that respect constraints on how data may be used. [事实] She says pure data-science backgrounds may lead people to focus on stronger correlations without realizing some risk dimensions cannot be used in pricing. [事实] She says A/B testing may be relatively unconstrained in marketing, but core actuarial work such as pricing and reserving is not. [推测] Data scientists entering insurance need to learn the regulatory and business environment before their models can be operationally useful.

[39:11] InsurTech and Subject-Matter Expertise

[事实] Mary Pat says InsurTech is real and already has successes, including use of AI in insurance. [事实] She says many successful efforts come from people with insurance background and practical knowledge. [事实] The host says data scientists need help from subject-matter experts such as actuaries to work successfully in insurance. [事实] Mary Pat says actuaries also need to know what models are available, what they can do, and how they can fail. [推测] The episode presents InsurTech success as depending less on algorithms alone and more on combining algorithms with insurance expertise.

[40:36] When Statistical Results Are Not Useful

[事实] Mary Pat says she has heard negative stories about statistics PhDs producing results that were not meaningful in an insurance environment. [事实] She gives an example where a model effectively implied that a company should go back in time and write more business in a profitable policy year. [事实] She says some recommendations are not feasible, such as selling more to one protected or constrained category. [推测] Model outputs need actionability checks, not just statistical validation.

[41:37] Advice for Aspiring Actuaries and Data Scientists

[事实] Mary Pat advises aspiring actuaries to balance exam study with attention to the actual business around them. [事实] She says she entered with strong math skills, but learning the business and how people behave was the harder part. [事实] The host says learning R or Python alone does not make someone a data scientist. [事实] Both speakers agree that practitioners need mathematics, statistics, coding, and understanding of the domain where the tools are applied. [推测] The career advice applies broadly: technical skill is necessary, but not sufficient without context and judgment.

[43:42] Closing Remarks

[事实] The host says viewers in and outside insurance can benefit from the conversation. [事实] He recommends following Mary Pat Campbell’s blog or online work for more information about actuarial science, mathematics, and statistics. [事实] He concludes that actuaries are key subject-matter experts in insurance and that data science should work with that expertise. [事实] The episode ends with thanks to Mary Pat and a request for viewers to like and subscribe. [推测] The closing reinforces the episode’s main message: data science and actuarial science are complementary when connected to real insurance practice.

播客点评/总结

This episode is valuable because it explains actuarial science through concrete data problems rather than abstract definitions. The strongest parts are the examples: COVID reporting lags, mortality tables, underwriting fraud, electronic health record errors, coding flags, impossible medical values, and currency mismatches.

The discussion is especially useful for data scientists or analysts who want to work in insurance. It makes clear that statistical strength is not enough; models also need regulatory awareness, domain knowledge, and operational usefulness.

[推测] The episode is less polished as a structured interview because the transcript includes repeated phrasing, informal detours, and a sponsor break in the middle. Still, the conversational style helps surface practical lessons that may not appear in a formal lecture.

[推测] Best suited listeners are aspiring actuaries, insurance data professionals, data scientists moving into InsurTech, and anyone who wants to understand why actuarial data work requires both quantitative methods and business judgment.