Giambattista Bloisi
|
d80f12da06
|
Build with spark 3.4 (dedup and dependencies only tested)
|
2023-07-10 15:54:48 +02:00 |
Giambattista Bloisi
|
861c368e65
|
Code for testing other grouping strategies
|
2023-07-10 15:52:35 +02:00 |
Giambattista Bloisi
|
dcc08cc512
|
Use UDAF and Aggregation class for testing
|
2023-07-07 12:35:30 +02:00 |
Giambattista Bloisi
|
df19548c56
|
small changes
|
2023-07-04 18:36:58 +02:00 |
Sandro La Bruzzo
|
890b49fb5d
|
optimized some dedup functions
|
2023-06-29 14:08:58 +02:00 |
Giambattista Bloisi
|
cb7ad9889c
|
Fix maven dependencies warning while building
|
2023-06-28 14:01:04 +02:00 |
Claudio Atzori
|
75ff902f9d
|
WIP: various refactors
|
2023-06-28 14:00:54 +02:00 |
Claudio Atzori
|
326367eccc
|
WIP: various refactors
|
2023-06-28 14:00:22 +02:00 |
Claudio Atzori
|
521dd7f167
|
WIP: various refactors
|
2023-06-28 14:00:18 +02:00 |
Claudio Atzori
|
649679de8d
|
WIP: various refactors
|
2023-06-28 13:59:11 +02:00 |
Sandro La Bruzzo
|
4c2dfcbdf7
|
Added first implementation using UDF function
|
2023-06-28 13:58:01 +02:00 |
Sandro La Bruzzo
|
9963fd6d29
|
updated log to add subentity
|
2023-06-28 13:36:05 +02:00 |
Sandro La Bruzzo
|
ed7e2ab6d1
|
reverted mistake on commit workflow.xml
|
2023-06-28 11:40:19 +02:00 |
Sandro La Bruzzo
|
9910ce06ae
|
added to CreateSimRel the feature to write time log
|
2023-06-28 11:38:16 +02:00 |
Sandro La Bruzzo
|
bd17c3edc8
|
added to CreateSimRel the feature to write time log
|
2023-06-28 11:20:58 +02:00 |
Claudio Atzori
|
909729a2fc
|
[dedup] tweaking num partitions, minor changes
|
2023-05-17 10:16:22 +02:00 |
Claudio Atzori
|
062abfd669
|
fixed NPE, removed unused stuff
|
2022-12-06 12:04:00 +01:00 |
Claudio Atzori
|
0aa725083f
|
extended dedup testing
|
2022-11-17 16:13:43 +01:00 |
Claudio Atzori
|
3dbc637d3e
|
code formatting
|
2022-11-17 09:55:41 +01:00 |
Claudio Atzori
|
ddff0e8999
|
merging duplicates using IdentifierComparator
|
2022-11-11 16:10:25 +01:00 |
Claudio Atzori
|
5af5a8ae42
|
added IdentifierComparator
|
2022-11-09 14:20:59 +01:00 |
Claudio Atzori
|
c26222623f
|
[maven-release-plugin] prepare for next development iteration
|
2022-04-07 13:32:22 +02:00 |
Claudio Atzori
|
86585a6b27
|
[maven-release-plugin] prepare release dhp-1.2.4
|
2022-04-07 13:32:19 +02:00 |
Claudio Atzori
|
ad85d88eaf
|
[maven-release-plugin] rollback the release of dhp-1.2.4
|
2022-04-07 13:28:35 +02:00 |
Claudio Atzori
|
598e11dfd7
|
[maven-release-plugin] prepare for next development iteration
|
2022-04-07 13:27:02 +02:00 |
Claudio Atzori
|
db3d9877a5
|
[maven-release-plugin] prepare release dhp-1.2.4
|
2022-04-07 13:26:58 +02:00 |
Claudio Atzori
|
3bba6d6e38
|
[maven-release-plugin] rollback the release of dhp-1.2.4
|
2022-04-07 12:23:17 +02:00 |
Claudio Atzori
|
2ac2d928bd
|
[maven-release-plugin] prepare for next development iteration
|
2022-04-07 12:18:47 +02:00 |
Claudio Atzori
|
85bc722ff4
|
[maven-release-plugin] prepare release dhp-1.2.4
|
2022-04-07 12:18:43 +02:00 |
Claudio Atzori
|
bc05b6168a
|
[maven-release-plugin] rollback the release of dhp-1.2.4
|
2022-04-07 11:49:06 +02:00 |
Claudio Atzori
|
505420fd61
|
[maven-release-plugin] prepare for next development iteration
|
2022-04-07 11:34:06 +02:00 |
Claudio Atzori
|
66e718981e
|
[maven-release-plugin] prepare release dhp-1.2.4
|
2022-04-07 11:34:02 +02:00 |
Claudio Atzori
|
61319b2e83
|
updated dhp-schema version; set entity-level dataInfo before & after merging the fields from the group of duplicates
|
2022-03-25 16:38:33 +01:00 |
miconis
|
c959639bd5
|
dependency updated to the new pace-core version
|
2022-03-15 16:33:03 +01:00 |
miconis
|
8991d097b4
|
bug fix in the DedupRecordFactory, DataInfo set before merge
|
2022-02-24 17:13:12 +01:00 |
Claudio Atzori
|
391aa1373b
|
added unit test
|
2022-01-19 17:13:21 +01:00 |
Claudio Atzori
|
44a937f4ed
|
factored out entity grouping implementation, extended to consider results from delegated authorities rather than identical records from other sources
|
2022-01-19 12:24:52 +01:00 |
Claudio Atzori
|
f4538f3c4c
|
cleanup
|
2021-11-19 11:33:10 +01:00 |
Claudio Atzori
|
2b46b87f56
|
fixed filtering criteria applied in SparkCopyRelationsNoOpenorgs to keep the parent/child relations from OpenOrgs
|
2021-11-19 11:30:29 +01:00 |
Claudio Atzori
|
a24b9f8268
|
[dedup] trivial refactoring
|
2021-11-18 17:12:02 +01:00 |
Claudio Atzori
|
c0750fb17c
|
avoid non necessary count operations over large spark datasets
|
2021-11-18 17:11:31 +01:00 |
Claudio Atzori
|
0a727d325d
|
[dedup] increased number of partitions in the consistency phase
|
2021-11-16 08:43:41 +01:00 |
miconis
|
611ca511db
|
set configuration property in openorgs duplicates wf
|
2021-10-07 15:39:55 +02:00 |
miconis
|
9646b9fd98
|
implementation of the http call for the update of openorgs suggestions
|
2021-10-07 11:29:11 +02:00 |
miconis
|
853333bdde
|
implementation of the whitelist for similarity relations
|
2021-09-20 16:21:47 +02:00 |
Claudio Atzori
|
9f4db73f30
|
updated/fixed unit tests
|
2021-08-11 15:02:51 +02:00 |
Claudio Atzori
|
2ee21da43b
|
suggestions from SonarLint
|
2021-08-11 12:13:22 +02:00 |
Claudio Atzori
|
2fff24df55
|
code formatting
|
2021-07-28 11:34:19 +02:00 |
Sandro La Bruzzo
|
3920c69bc8
|
change implementation of resolve Relation to generate jsonRdd in output
|
2021-07-25 09:51:36 +02:00 |
Sandro La Bruzzo
|
058b636d4d
|
added control to check if the entity exists
|
2021-07-22 16:08:54 +02:00 |