Week 7
Unstructured Data
- Instructors
- Maclean Gaulin
What is unstructured data?
- Non-tabular data
- Text
- Images
- Video / audio
- Generally any data without pre-defined structure
- Most data are unstructured
- Files from client often unstructured
Why do we care?
- Majority of data are unstructured (80-90% estimates)
- Communication
- Records & evidence (e.g. receipts, invoices)
- Regulations & guidance
- Wealth of information, but only if translated to something computers and algorithms can use
- ML methods increasingly capable of conversion
What to do with Unstructured Data?
- You can’t run a regression on unstructured data
- Must add “structure” to make data useful
- Could be conversion to tabular form
- Could be analysis
- Feature extraction is term for creating variables (features) from data
- Covers gamut from word count to LLM parsing
OCR
- This is a receipt from my recent shopping trip. Please transcribe it into a table.
| Purchased Items | |
|---|---|
| LP HOTPOT SAUCE | $4.39 FTO |
| Disc. $2.00 | |
| LKK FRIED TOFU BALL | $2.99 FTO |
| ASN SOY BEAN PASTE | $3.59 FTO |
| 1 @ 3.39/lb $3.39/lb | |
| LOTUS ROOT | $3.90 FY1 |
| MEME | |
| XFY SPICY SEASONING BASE | $4.99 FTO |
| XFY MUSHROOM SOUP BASE | $4.99 FTO |
| XFY SPICY SEASONING BASE | $4.99 FTO |
| LP HOTPOT BASE | $4.99 FTO |
| Disc. $2.00 | |
| LP HOTPOT SAUCE | $4.99 FTO |
| Disc. $2.00 | |
| LP HOTPOT BASE | $4.99 FTO |
| Disc. $2.00 | |
| SJX HOT POT BASE | $13.99 FTO |
| Summary | |
| Item Count: | 11 |
| Subtotal: | $59.40 |
| Tax: | $1.79 |
| Total: | $61.19 |
| Payment: | $61.19 |
| Cash Change: | $0.00 |
| You Saved: | $8.00 |
| Payment Method | |
| Type: | CREDIT SALE |
| Card: | VISA ***1190 |
| Entry Method: | CONTACTLESS |
| Auth. Code: | 64211 |
| Amount: | $61.19 |
- Gemini prompt: This is a receipt from my recent shopping trip. Please transcribe it into a table.
- Output (including table to right):
- Store Information:
- CHINATOWN SUPERMARKET
- 3390 South State Street
- Salt Lake City, UT 84115
- (801) 906-8788
- Transaction Details:
- Station 04 2041752958
- 05/19/2025 12:06:19
- Cashier #: 275
- AUTH CODE: 069611
- ENTRY METHOD: CONTACTLESS
- ACCT: VISA 1130
Image Analysis Caveats
- Only as good as the algorithm
- Off the shelf solution may have noise or bias (based on the assumptions is makes or features it has)
- Customization is costly, should be balanced against errors (type and amount)
What are features?
- Usable information, measures, signals
- Summarization / aggregation
- Extraction / filtering
- We want to choose features that capture desired info
- Tone of investor call
- Measuring inventory level estimates from images
- Extracting text from image
Common text features
- Simple: text length, word counts
- Sometimes simple is good enough
- Intermediate: tone, sentiment, topics
- Broad generalizations / aggregation from the text
- Advanced: specific information
- Targeted to extract specific facts
- Information retrieval subfield
- LLMs are now quite good at this
Natural Language Processing
- NLP addresses understanding language to process and analyze texts
- E.g., topic analysis, information retrieval
- Old approach: developed rules (grammars) for parsing text
- E.g., regular expressions
- New approach: lots of math (machine learning)
NLP Pipeline
- Preprocessing
- clean text
- tokenize
- stemming
- remove/replace words
- Feature Engineering
- Part of speech tags
- Entity parsing (NER)
- Quantize
- N-gram generation
- Vectorization
- (e.g. BOW)
- Weighting
- (e.g. TF-IDF)
- Analysis
- Topic Analysis (clustering)
- Sentiment Analysis
- (classification)
- Summarization
- Information Retrieval
Preprocessing
- Tokenization: breaking text into tokens (letters, words, sentences, or whatever is needed)
- “University of Utah” [“University”, “of”, “Utah]
- Cleaning: lowercase, remove punctuation, add placeholders (e.g. numbers, dates)
- “FYEnd, 12/31/2025” [“fyend”, “DATE”]
- Stemming: replacing words with their “root” by removing conjugations and plurality
- “She abhors assumptions” [“she”, “abhor”, “assum”]
Preprocessing Caveats
- Preprocessing removes information
- Uppercase letters can be informative (May vs may)
- Punctuation is often important
- Conjugation is almost always important
- When using for analytics, preprocessing is critical for the signal-to-noise tradeoff
- Removing noise if case, punctuation, conjugation etc. aren’t important
- Removing information if they are
Vectorization
- Turning text into numbers
- “See spot Run. Run Spot run”
- [see, 2 spot, 3 run]
- <1, 2, 3, 0, …, 0>
- What if we want to keep word the sense of words being together (i.e. word order): bi-grams
- [“see spot”, “spot run”, “run spot”, “spot run”] (“run run”?)
- <1, 2, 1, 0, …, 0>
N-Grams
- Vector of words that are N-tokens long
- Many more bi-grams than unigrams
- From my paper on Executive Compensation:
- # of 1-grams: 238,001
- # of 2-grams: 2,230,332 (9.4x)
- # of 3-grams: 3,459,221 (1.55x)
- # of 4-grams: 5,386,503 (1.56x)
NLP in Accounting
- Information retrieval: extracting specific information or facts from documents
- Entity linking: identifying people/places/things and linking them to known entities
- Relationship extraction: extracting relationships between things or ideas
- Ex: company product, event date
- Topic analysis & sentiment analysis
Sentiment Analysis
- Classifies text into positive, negative, neutral
- Old method: rule or dictionary based
- Look for negative or positive words
- New method: statistically based
- Requires examples for training
- Off the shelf generic LLM solution
Sentiment in Accounting
- Analyzing competitors' disclosures
- Reviewing investor sentiment on conference calls
- Demand estimation from customer responses/posts
- Identifying early risk indicators from social media
Topic Modeling
- Categorize content of text into topics
- Customized to content and desired topics
- Models:
- LDA: define # of topics, needs lots of texts
- NMF: better for shorter texts and less data
- BERTopic: more advanced, single topic per text
Topic Modeling Example – Risk Factors
LLMs for NLP
- It all comes down to the prompt (for now)
- Very good at:
- Named entity recognition, relation extraction, sentiment
- Pretty good at:
- Entity linking, fact retrieval
- Unknown good at:
- Topic analysis, drawing conclusions
NLP in the LLM Era
- LLMs may be overkill
- LLMs are far more black-box
- Dependent on training data, fine-tuning data, but to an unknown extent
- Still room for older approaches
- When it does what it does well
- When a (more) deterministic outcome is necessary
What about images?
- Increasingly used in accounting
- Images of receipts for reimbursement
- Optical character recognition (OCR)
- Inventory tracking and verification
- Evidence for asset write-offs
- Important to understand caveats and limitations of image processing and use
Image processing basics
- Image processing used to look like historical NLP, lots of steps, transforms, and heuristics
- Line detection, distance/area calculations, thresholding, segmentation
- Image analysis tasks were very customized and expensive to design, develop, and deploy
Example – counting Sheep
- Load image
Example – counting Sheep
- Load image
- Make greyscale
Example – counting Sheep
- Load image
- Make greyscale
- Threshold to BW
Example – counting Sheep
- Load image
- Make greyscale
- Threshold to BW
- Remove small objects
Example – counting Sheep
- Load image
- Make greyscale
- Threshold to BW
- Remove small objects
- Shrink objects
Example – counting Sheep
- Load image
- Make greyscale
- Threshold to BW
- Remove small objects
- Shrink objects
- Count remaining
- = 3,076 (actually 3,067)
Gemini – Count the Sheep
- Prompt: This is an image of sheep taken from a drone. Please count the sheep in this image.
- Answer: I am unable to provide an exact count of the sheep in the image. Due to the high number of sheep, their close proximity, and the aerial perspective, it is very difficult to distinguish and count each individual animal accurately.
- Attempting a manual count would be prone to significant error. Automated counting through image analysis would require specialized software and would still likely face challenges with overlap and resolution.
- Therefore, I cannot fulfill your request for an exact count of the sheep in this image.
ChatGPT – Count the Sheep
- Prompt: This is an image of sheep taken from a drone. Please count the sheep in this image.
- Answer: Counting the exact number of sheep in this image is challenging due to their dense and overlapping formation. However, I can use image processing techniques to estimate the number. Let me process the image to give you a sheep count estimate.
- [code]
- There are approximately 2,347 sheep in the image based on contour detection from the drone photo. Here is the visual overlay showing the detected sheep outlined in green.
- actually 3,067
Image processing Now
- Machine learning introduced statistical methods to segment and identify objects in images
- Object detection, classification, etc.
- Off the shelf solutions for many tasks
- The EY sheep case used countthings.com
- Customizing and fine-tuning often necessary
Regular Expressions
- Regular expression (regex) is a search pattern
- Like a “find” but with flexibility
- Common uses:
- Search, optionally replace (e.g. find all acronyms)
- Verify data (e.g. phone number is correctly entered)
- Extract data (e.g. extract year from a long date)
Using Regex
- Write a “regex”, which is the string to “search”
- Literal:
invoice,2024 - Special characters:
[JFMASOND]\w+ 202\d - Use that regex to search through data
- Requires software to “apply” the regex, e.g. Python
Regex Basics
- Literals:
- Letters:
A-Z - Numbers:
0-9 - Special characters (many need escaping):
<>(-,.!) - Character Classes
-
\d= any digit (0-9) -
\w= word character (A-Z, a-z, 0-9, _) -
\s= whitespace (space, tab) -
[]= any of the characters in brackets (\d=[0-9])
Examples
- Pattern:
[A-Z]\d\d\d - Matches:
A123,B456,Z999,ABC123 - Doesn't Match:
1234,AB12,a123 - Pattern:
[JFMASOND]\w\w 202\d - Matches:
Jan 2020,MAY 2029,All 2025 - Doesn't Match:
jan2020,Oct. 2025
Special Characters – Repeats
| Symbol | Definition | Example |
|---|---|---|
| * | Zero or more | \d* matches “”, “1”, “123” |
| + | One or more | \d+ matches “1”, “123”, not “” |
| ? | Zero or one (optional) | colou?r matches color, colour |
| {n} | Exactly n times | \d{4} matches “1234”, “2024” |
| {n,m} | Between n and m times | \d{2,4} matches “24”, “2024” |
Data Extraction – Capture Groups
- Parentheses
()define “groups” - Useful for using
|, the “or” character -
(January|February|…|December) matches 1 month - Groups can also “capture” data, usually for extracting
-
INV-(\d{7}) 7 digit number only, no INV- - Doesn’t change the search, tells regex what to return
Lab 7
- Regular Expressions
- Clean dataset, add new usable columns
- Submit your regex expressions
- Submit your cleaned dataset
Project 2
- Due Sunday, February 22
- Project 1: FSA ratio(s) in cross section & over time
- Project 2: Returns in cross section (FSA ratio, P1 cross section) and over time
- You can choose one, multiple, interactions, etc.
- I will be looking for continuity between your RQ and economic story and your tests
- I.e., how you calculate returns should match the economics